llama-server is the HTTP server that ships with llama.cpp: it loads a GGUF model and serves an OpenAI-compatible API plus a browser chat UI, and you start it with one command, llama-server -hf ggml-org/gemma-4-E4B-it-GGUF:Q4_0 (or -m model.gguf for a local file). It listens on http://127.0.0.1:8080 by default, so the chat UI and the /v1/chat/completions endpoint are both at that address once the model loads (Source: llama.cpp server README).
The current release also exposes it as llama serve, with identical flags (Source: llama.app serve docs). We ran every command below on the stable llama.cpp v0.5.0 release, build b11146 (commit 7fe450e19), on an Apple M1 with 8 GB of RAM, using the official macOS arm64 download and a 0.5B Qwen model (Source: llama.cpp v0.5.0 release).
Key takeaways
- Start it:
llama-server -m model.gguforllama-server -hf <user>/<repo>:<quant>; the default quant for-hfisQ4_K_M. - Defaults to know: host
127.0.0.1, port8080, context from the model, GPU layersauto, slotsauto, web UI on,/metricsoff. - Secure it: add
--api-keybefore you bind0.0.0.0;/healthand the web UI page stay public either way. - Connect clients: point any OpenAI SDK or Open WebUI at
http://host:8080/v1.
Install llama-server and check the version
llama-server comes inside every llama.cpp release archive, so there is nothing separate to install. The GitHub release page ships a macOS arm64 tarball, Linux builds for CPU, CUDA, ROCm and Vulkan, and Windows zips per GPU backend, each with llama-server inside (Source: llama.cpp v0.5.0 release).
Stable releases now use semantic versions, and each points at a nightly build tag: v0.5.0 maps to b11146 (Source: llama.cpp v0.5.0 release).
Measured in our test: the llama-b11146-bin-macos-arm64.tar.gz asset was 10.7 MB, unpacked to 27 MB, and contained 22 llama* executables. ./llama-server --version printed version: 0.5.0-dev (build 11146, commit 7fe450e19), and ./llama --help listed serve cli download version licenses help.
On macOS, run xattr -dr com.apple.quarantine . in the unpacked folder before the first start. On Windows the binary is llama-server.exe in the zip for your GPU, with the same flags; our llama.cpp Windows guide covers which zip to pick. If a model does not fit, choose a smaller quant with our Q4_K_M vs Q5_K_M comparison.
Start llama-server: the one command and a real one
The minimal start needs only a model. -m takes a local GGUF path, while -hf downloads from Hugging Face and caches the file, with an optional :quant suffix (Source: llama.cpp server README).
# local file
llama-server -m ./qwen2.5-0.5b-instruct-q4_k_m.gguf
# from Hugging Face (quant defaults to Q4_K_M)
llama-server -hf ggml-org/gemma-4-E4B-it-GGUF:Q4_0For a server you keep running, set context, GPU offload, slots, an alias and a key explicitly. This is the command we tested, key redacted:
llama-server -m qwen2.5-0.5b-instruct-q4_k_m.gguf -a qwen2.5-0.5b \
-c 4096 -ngl all -np 2 \
--host 127.0.0.1 --port 8089 \
--api-key "$LLAMA_API_KEY" --jinja \
-ctk q8_0 -ctv q8_0 --metricsMeasured in our test: that command reached a 200 on /health in 793 ms, the process held 610 MB resident, and the server log printed listening on http://127.0.0.1:8089. The Qwen model decoded at a median 94.2 tokens per second over five runs with every layer on the M1 GPU.
Launched with no model at all, llama serve starts router mode, which loads and unloads models from the Hugging Face cache, a --models-dir folder or a --models-preset file on demand (Source: llama.app serve docs).
llama-server flags reference
Every option can also be set through an environment variable, which is how Docker setups configure the server; a command-line flag wins when both are set (Source: llama.cpp server README).
Measured in our test: llama-server --help in build 11146 printed 731 lines with 259 option rows, and 148 of those options list an env: variable. The table covers the flags almost every setup touches, with defaults copied from that output.
| Flag | What it does | Default in b11146 | Env variable |
|---|---|---|---|
-m, --model FNAME | Load a local GGUF file | none | LLAMA_ARG_MODEL |
-hf, --hf-repo user/model[:quant] | Download and cache from Hugging Face | quant Q4_K_M | LLAMA_ARG_HF_REPO |
-c, --ctx-size N | Total context in tokens | 0 = from model | LLAMA_ARG_CTX_SIZE |
-ngl, --gpu-layers N | Layers offloaded to GPU: number, auto or all | auto | LLAMA_ARG_N_GPU_LAYERS |
-np, --parallel N | Server slots for concurrent requests | -1 = auto | LLAMA_ARG_N_PARALLEL |
--host HOST | Bind address; comma list or .sock path | 127.0.0.1 | LLAMA_ARG_HOST |
--port PORT | Listen port | 8080 | LLAMA_ARG_PORT |
--api-key KEY | Require a key; comma list allowed | none | LLAMA_API_KEY |
--api-key-file FNAME | One key per line | none | LLAMA_ARG_API_KEY_FILE |
-a, --alias NAME | Model name the API reports | file path | LLAMA_ARG_ALIAS |
--jinja / --no-jinja | Jinja chat templates, needed for tool calls | enabled | LLAMA_ARG_JINJA |
-ctk / -ctv TYPE | KV cache type for K and V (f16, q8_0, q4_0 and more) | f16 | LLAMA_ARG_CACHE_TYPE_K |
-fa, --flash-attn | Flash attention on, off, auto | auto | LLAMA_ARG_FLASH_ATTN |
--ui / --no-ui | Built-in web UI | enabled | LLAMA_ARG_UI |
--metrics | Prometheus /metrics endpoint | disabled | LLAMA_ARG_ENDPOINT_METRICS |
--slots / --no-slots | /slots monitoring endpoint | enabled | LLAMA_ARG_ENDPOINT_SLOTS |
-to, --timeout N | Read/write timeout in seconds | 3600 | LLAMA_ARG_TIMEOUT |
Two defaults changed from older guides. --jinja is now on unless you pass --no-jinja, and -ngl defaults to auto with a --fit pass that trims unset options to fit device memory (Source: llama.cpp server README). The KV cache type trades memory for accuracy; our KV cache size guide shows how many gigabytes q8_0 saves at long context.
Parallel slots and context size
A slot is one in-flight conversation, and -np sets how many the server keeps. The llama.app docs describe the context as shared across slots, but the split you get depends on whether the KV cache is unified (Source: llama.app serve docs).
With an explicit -c 4096 -np 2, the server divides the context evenly between the two slots.
Measured in our test: -np 2 produced n_slots = 2, n_ctx_slot = 2048, kv_unified = 'false', and /slots reported 2048 tokens for each of the two slots. Two simultaneous /completion requests came back with id_slot values [0, 1], so they ran side by side.
Leaving both -c and -np unset behaved differently. Measured in our test: llama serve -m model.gguf chose n_slots = 4, n_ctx_slot = 32768, kv_unified = 'true', the model's full trained context of 32768 tokens, with one unified pool that the four slots draw from.
Inference: if a client reports less context than the model supports, check -np. Four explicit slots on -c 32768 leave each conversation 8,192 tokens.
Endpoints: the OpenAI-compatible API and native routes
llama-server exposes three API families on one port: OpenAI-compatible routes under /v1, an Anthropic-compatible /v1/messages, and native llama.cpp routes with extra options (Source: llama.app API docs).
| Endpoint | Purpose | No key (server has --api-key) | With key |
|---|---|---|---|
GET /health, /v1/health | Liveness: 200 ready, 503 loading | 200 | 200 |
GET / | Web UI page | 200 | 200 |
GET /v1/models | Model id, trained context | 401 | 200 |
POST /v1/chat/completions | OpenAI chat, streaming, tools | 401 | 200 |
POST /v1/messages | Anthropic Messages (uses x-api-key) | 401 | 200 |
POST /completion | Native completion with id_slot, timings | 401 | 200 |
GET /props | Build info, slots, chat template | 401 | 200 |
GET /slots | Per-slot state | 401 | 200, or 501 with --no-slots |
GET /metrics | Prometheus counters | 401 | 200 with --metrics, else 501 |
Every status code above came from our run; the purpose column comes from the docs. The README also lists /v1/completions, /v1/responses, /v1/embeddings, /v1/rerank and /infill (Source: llama.cpp server README).
Measured in our test: a temperature-0 chat request returned Paris with system_fingerprint b11146-7fe450e19, 38 prompt tokens, and a llama.cpp timings object alongside the standard OpenAI usage block. /metrics exposed 15 series such as llamacpp:prompt_tokens_total, and without -a the model id was the full file path.
The built-in web UI
Open http://127.0.0.1:8080 in a browser and llama-server serves its own chat interface, enabled by default and switched off with --no-ui (Source: llama.app serve docs). It handles plain chat, and with a multimodal model you can drag images into the conversation.
The page itself does not ask for the API key. Measured in our test: GET / returned 200 with text/html and no key even though --api-key was set, while every /v1 call without the key returned 401. A plain curl of / without Accept-Encoding: gzip returned Error: gzip is not supported by this browser, so use a real browser or curl --compressed when you check it.
Inference: since the page is public, anyone who reaches the port learns a llama.cpp server runs there, even without chat access.
Tool calling with --jinja
Tool calling works through the normal OpenAI tools and tool_choice fields, and llama.cpp parses the model's native format for families trained on tool use, with a generic fallback for the rest (Source: llama.app API docs). It needs the model's Jinja chat template, which is why older guides tell you to add --jinja; in b11146 it is already the default.
Measured in our test: with a get_weather tool defined, the small Qwen model returned finish_reason tool_calls and the call {"name": "get_weather", "arguments": "{\"city\": \"Paris\"}"}.
Keep --tools separate from this. That flag turns on the server's built-in agent tools, including exec_shell_command and write_file, and the README warns not to enable it in untrusted environments (Source: llama.cpp server README).
Run llama-server in Docker
The official image is ghcr.io/ggml-org/llama.cpp:server, which contains only the llama-server binary; GPU builds are tagged server-cuda (CUDA 12), server-cuda13, server-rocm, server-vulkan and server-intel (Source: llama.cpp Docker docs).
docker run -p 127.0.0.1:8080:8080 -v ./models:/models \
ghcr.io/ggml-org/llama.cpp:server \
-m /models/model.gguf -c 4096 \
--host 0.0.0.0 --port 8080 --api-key "$LLAMA_API_KEY"Inference: --host 0.0.0.0 is needed inside the container so the published port reaches the server, and publishing as 127.0.0.1:8080:8080 keeps it on the host's loopback rather than the LAN. On NVIDIA hosts add --gpus all and the server-cuda tag. We did not run the container on the M1; image names and flags come from the docs.
For Compose, every flag has an LLAMA_ARG_* form, and the README example sets LLAMA_ARG_MODEL, LLAMA_ARG_CTX_SIZE and LLAMA_ARG_N_PARALLEL (Source: llama.cpp server README).
Connect Open WebUI and the OpenAI SDK
Any OpenAI client works once you change the base URL to http://127.0.0.1:8080/v1 and pass the key as the API key (Source: llama.app API docs).
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8080/v1", api_key="YOUR_KEY")
r = client.chat.completions.create(
model="qwen2.5-0.5b",
messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)Measured in our test: the openai Python package 3.22.1 completed a chat round trip against our server with only base_url and api_key changed.
In Open WebUI, go to Settings, Admin, Connections, OpenAI and add a connection. Use http://127.0.0.1:8080/v1, or http://host.docker.internal:8080/v1 when Open WebUI runs in Docker, and keep the /v1 on the end (Source: Open WebUI llama.cpp guide). Enter your key under Bearer auth, and set Provider to llama.cpp to get the loaded-model indicator and the Eject button. For other front-ends, our Ollama GUI roundup covers clients that also accept an OpenAI-compatible URL.
Secure llama-server: API key and 0.0.0.0
The default bind is 127.0.0.1, so a fresh server is reachable only from the same machine (Source: llama.cpp server README).
Measured in our test: both servers listened on 127.0.0.1 only according to lsof, and a second server started without --api-key answered /v1/models with 200 to anyone who could reach it.
Passing --host 0.0.0.0 changes that. Copy-paste examples, including the llama.app Compose file, bind 0.0.0.0 without a key, which leaves an open inference endpoint on your network (Source: llama.app serve docs). The README's own CORS guidance for a public server is to set an API key and put the server behind a reverse proxy (Source: llama.cpp server README).
A safe default, in order:
- Keep
--host 127.0.0.1unless another machine needs access. - Set
--api-keyor--api-key-filebefore binding any other address; a wrong key returns 401. - Terminate TLS at a reverse proxy, or build with SSL and use
--ssl-key-fileand--ssl-cert-file. - Avoid
--toolsand--agenton any shared host. - Remember
/healthand the web UI page stay public by design.
FAQ
What is a LLaMA server?
A LLaMA server usually means llama-server, the HTTP server bundled with llama.cpp. It loads a GGUF model on your own machine and serves an OpenAI-compatible API plus a web chat UI on port 8080. Despite the name, it runs most open model families, including Qwen, Gemma and Mistral, not only Meta's Llama.
How do I run the LLaMA server?
Download a llama.cpp release for your OS, unpack it, and run llama-server -m model.gguf or llama-server -hf user/repo:quant. Newer releases also accept llama serve with the same flags. Wait for the log line listening on http://127.0.0.1:8080, then open that address in a browser or call /v1/chat/completions.
What is LLaMA used for?
Llama is Meta's family of open-weight language models, used for chat assistants, coding help, summarization, retrieval-augmented search and agents. With llama-server you run Llama or any other GGUF model locally, which keeps data on your hardware and lets existing OpenAI-based apps switch to a local model by changing one base URL.
What are the differences between LLaMA.cpp and LLaMA-server?
llama.cpp is the whole project: the C/C++ inference library plus a set of tools. llama-server is one of those tools, the HTTP server that wraps the library with an OpenAI-compatible API, parallel slots and a web UI. llama-cli is the terminal chat tool from the same release, sharing most model flags.
How do I set an API key for llama-server?
Start the server with --api-key YOUR_KEY, a comma-separated list for several keys, or --api-key-file keys.txt with one key per line. The LLAMA_API_KEY environment variable works too. Clients send Authorization: Bearer YOUR_KEY; requests without it get a 401, while /health and the web UI page stay public.
Related coverage
- Ollama vs llama.cpp: if you are still choosing between llama-server and Ollama's server, start with this comparison.
- Q4_K_M vs Q5_K_M vs Q8_0: which GGUF quant to pass to
-hf. - KV cache size: what
-c,-npand-ctk q8_0cost in memory. - Local LLM hardware calculator: check whether a model fits before you set
-ngl.
References
- llama.app API docs - https://llama.app/docs/api
- llama.app serve docs - https://llama.app/docs/serve
- llama.cpp Docker docs - https://github.com/ggml-org/llama.cpp/blob/master/docs/docker.md
- llama.cpp server README - https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md
- llama.cpp v0.5.0 release - https://github.com/ggml-org/llama.cpp/releases/tag/v0.5.0
- Open WebUI llama.cpp guide - https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-llama-cpp/

