llama.cpp
Atomic supports the llama.cpp router server. The router discovers multiple GGUF models and loads or unloads them on demand. Use a current llama.cpp build with router support. Follow its build instructions or install a prebuilt release.Start the router
Startllama-server without --model or -m; those options start single-model mode instead of router mode.
--models-dirdiscovers local GGUF files.--no-models-autoloadleaves loading under explicit/llamacontrol.--jinjaenables compatible chat templates and tool calling.-ngl 999offloads as many layers as possible to the GPU.-c 32768sets each model’s context window. Omit it to use the model’s native context, which may require substantially more memory.
Configure Atomic
Run:http://127.0.0.1:8080. The same values can be supplied without /login:
llama-server with the matching --api-key. Keep --host 127.0.0.1 for local-only access.
Manage models
Run/llama in interactive mode:
- Select an unloaded model to load it, or a loaded model to unload it.
- Select Download model…, search Hugging Face, then choose a repository and quantization. Exact
owner/repository[:quant]values also work. - Press Escape during a load or download to confirm cancellation.
HF_TOKEN when set, then checks $HF_TOKEN_PATH, $HF_HOME/token, $XDG_CACHE_HOME/huggingface/token, and ~/.cache/huggingface/token. Unauthenticated search has lower rate limits. Atomic warns before gated downloads and links to the access page. Because llama.cpp performs the download, its process must also have HF_TOKEN for gated repositories.
Atomic asks before unloading other models, never silently unloads models, and never deletes model files. /llama always displays the router’s current state because other clients may share it. Only loaded models appear in /model; load one first, then select it there. Atomic persists the last successful loaded-model catalog in ~/.atomic/agent/models-store.json (or the active custom agent directory), so those entries remain selectable after restart and before the first successful refresh. If that first online refresh fails, Atomic reports the original router error while continuing to expose the validated persisted catalog; a later successful refresh replaces those stale loaded-state entries without duplicates. If the router disconnects, choose Retry to reconnect and refresh state without replaying the interrupted operation.
Each llama model uses the router-reported loaded context (meta.n_ctx, then training context, otherwise Atomic’s fallback) for both contextWindow and maxTokens; Atomic no longer applies a separate 16K output cap. The server remains authoritative and may impose a smaller practical generation limit.
Troubleshooting
- No models in
/llama: Check--models-dir, the directory layout, and restart the router. - Model missing from
/model: Load it with/llamafirst. - Load fails or uses too much memory: Lower
-cor unload another model. - Server is not in router mode: Start it without
--model,-m, or-hf.