Quick answer
Ollama and llama.cpp are tools for local work: a laptop, a home server, a single workstation, a prototype, one user or a few people. They are very good at it. vLLM and SGLang are production engines: you put them behind a company API that dozens of people and agents use at the same time. You won't see the difference with one question at a time. You'll see it when a dozen people ask at once and an agent sends long prompts full of documents. We serve models with vLLM and SGLang, and we decide which of them runs a given model by measuring.
As of 20 August 2026: vLLM 0.27.1, SGLang 0.5.17, Ollama 0.32.15, and llama.cpp released as daily builds.
Two worlds: local and production
| Ollama | llama.cpp (llama-server) | vLLM | SGLang | |
|---|---|---|---|---|
| Built for | laptop, home, prototype, a few people | single user, any hardware | company API, many users and agents | company API, agents, long shared prefixes |
| Hardware | CPU, GPU, Mac | CPU, NVIDIA, AMD and Intel GPUs, Vulkan, Mac | NVIDIA, AMD and Intel GPUs, CPU, others via plugins | NVIDIA and AMD GPUs, Intel Xeon CPUs, TPUs, Ascend NPUs |
| Weight format | GGUF (Ollama library) | GGUF | Hugging Face safetensors (BF16, FP8, NVFP4, AWQ, GPTQ and more), also GGUF | Hugging Face safetensors, FP4, FP8, INT4, AWQ, GPTQ quantizations |
| Concurrent requests | 1 per model by default | slots, continuous batching on by default | continuous batching, PagedAttention | continuous batching, RadixAttention |
| Prompt reuse | not described in the FAQ | prompt cache per slot | automatic prefix caching | prefix tree (RadixAttention), optional cache in RAM and on disk (HiCache) |
| API | own + OpenAI-compatible | OpenAI-compatible | OpenAI-compatible, Anthropic Messages, gRPC | OpenAI-compatible, Anthropic Messages |
| License | MIT | MIT | Apache 2.0 | Apache 2.0 |
Sources for the table: Ollama FAQ, llama.cpp server README, vLLM README, SGLang README, SGLang: HiCache, SGLang: Anthropic-Compatible API.
Why does the number of users change everything?
A language model generates its answer token by token. For every token the GPU has to read all the model's weights from memory. When one person asks, the card spends most of its time waiting for memory and its compute sits idle. When a dozen people ask at once, the engine can process their requests in one batch: the weights are read once and a dozen answers are computed.
The catch is the KV cache, the memory that holds each conversation's context. It grows with prompt length and with the number of concurrent requests. How to work it out is shown in our guide how much VRAM for a local LLM.
Engines differ in how they manage that cache:
- Reserving up front. Each slot gets a fixed chunk of memory for the maximum context. Simple, but it wastes memory, because most conversations are shorter than the maximum. That's how local tools work, where one person is asking anyway.
- PagedAttention (vLLM). The KV cache is split into small blocks allocated on demand, like memory pages in an operating system. The vLLM authors showed this delivers 2-4 times the throughput at the same latency compared with earlier systems (Kwon et al., 2023).
- RadixAttention (SGLang). The KV cache of all requests is kept in a prefix tree (radix tree). A new request that starts the same way as an earlier one reuses the part already computed, and rarely used branches are evicted when memory runs short (Zheng et al., 2023). This suits agents well: they branch conversations and keep sending the same system prompt and the same tool definitions.
- Continuous batching. A new request joins the running batch as soon as there's room, instead of waiting for every answer in the batch to finish.
- Prefix caching. When many requests start the same way, the engine doesn't recompute the shared beginning. In vLLM this is automatic prefix caching (vLLM: Automatic Prefix Caching), in SGLang the tree described above. In RAG and agents this saves a lot of time to first token.
Ollama and llama.cpp: great locally
Ollama is the fastest way to run a model on your own computer: one command downloads the weights and starts a server. Under the hood it uses llama.cpp and the GGUF format. It has an OpenAI-compatible API, so an application written against Ollama can later move to a production engine without a rewrite (Ollama: OpenAI compatibility).
Why we don't put it behind a company API follows directly from the documentation (Ollama FAQ):
OLLAMA_NUM_PARALLELdefaults to 1, so one model handles one request at a time. Further requests go to a queue (up to 512 by default).- Raising parallelism increases the memory needed for context in proportion to the number of parallel requests and the context length.
- Ollama loads models on demand and keeps several in memory (up to 3 per GPU by default). On a workstation that's convenient; on a production server it means unexpected reloads.
llama.cpp is the library and llama-server that Ollama, among others, is built on. It runs on CPUs, NVIDIA (CUDA), AMD (HIP) and Intel (SYCL) cards, via Vulkan and on Macs (Metal) (llama.cpp README), and GGUF quantizations let a model fit on hardware with little memory, at a cost in quality at low bit widths. The server can do more than people usually assume: slots with an automatically chosen count, continuous batching on by default, a shared KV cache buffer, a prompt cache, OpenAI-compatible endpoints, an API key and a router mode that loads several models from a directory (server README).
What to use them for:
- a developer's laptop and a private assistant,
- a home server or a single workstation,
- learning, testing new models, prototypes,
- Macs, CPUs, non-NVIDIA GPUs, edge devices,
- models that need heavy quantization to fit at all.
What we don't trust them with: a company API used at the same time by a dozen or several dozen people and agents. That's where engines built around many concurrent conversations pull ahead.
vLLM and SGLang: production engines
When a model has to serve the whole company, meaning a dozen or several dozen people at once, agents and batch document processing, we pick one of two engines. Both are Apache 2.0 licensed and have a very similar feature set.
vLLM offers in one package (vLLM README):
- PagedAttention, continuous batching, chunked prefill and prefix caching,
- FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ and GGUF quantization,
- speculative decoding,
- tensor, pipeline, data and expert parallelism to split a model across several GPUs,
- many LoRA adapters on one base model (vLLM: LoRA),
- structured outputs plus tool-call and reasoning parsers for agents,
- an OpenAI-compatible API, plus Anthropic Messages API and gRPC support,
- support for NVIDIA, AMD and Intel GPUs, CPUs and other accelerators via plugins.
SGLang has a similar scope (SGLang README):
- RadixAttention, i.e. prefix caching in a tree, continuous batching, paged attention and chunked prefill,
- a hierarchical KV cache (HiCache): GPU, RAM and external storage, useful for long context and multi-turn conversations (SGLang: HiCache),
- FP4, FP8, INT4, AWQ and GPTQ quantization (SGLang: Quantization),
- speculative decoding and prefill-decode disaggregation onto separate machines,
- tensor, pipeline, expert and data parallelism,
- many LoRA adapters in one batch, based on S-LoRA and Punica techniques (SGLang: LoRA Serving),
- structured outputs, tool and reasoning parsers (SGLang: Tool Parser),
- an OpenAI-compatible API and an Anthropic-compatible
/v1/messagesendpoint, enabled by default (SGLang: Anthropic-Compatible API).
How do they differ in practice?
| vLLM | SGLang | |
|---|---|---|
| KV cache management | PagedAttention, automatic prefix caching | RadixAttention (prefix tree), HiCache |
| Strong point | very broad model and hardware support, frequent releases | workloads with long shared prefixes: agents, multi-turn conversations |
| Configuration | tool and reasoning parser chosen per model family | tool and reasoning parser chosen per model family |
Both projects move very fast and borrow ideas from each other, so feature differences blur from release to release. Performance differences depend on the model, the quantization, prompt lengths and the engine version. That's why we don't pick an engine from a ranking; we measure both on the traffic the model will serve.
The price of this power, for both: more configuration than Ollama, weights in Hugging Face format, frequent releases (a new version every few weeks), and new models often needing the latest engine version.
What about TGI?
Since December 2025, Hugging Face's Text Generation Inference has been in maintenance mode: it only accepts minor fixes and documentation. Hugging Face itself recommends vLLM, SGLang and local engines such as llama.cpp and MLX (TGI README). The last release, 3.3.7, dates from December 2025. It's not worth building new deployments on it.
Weight formats: GGUF or safetensors?
A practical difference that's easy to miss:
- GGUF is the format of llama.cpp and Ollama. It has many quantization variants (e.g. Q4, Q5, Q8) and works well on CPUs and Macs.
- Hugging Face safetensors is the format in which vendors publish models and official quantizations (FP8, NVFP4, AWQ). vLLM and SGLang work with it.
vLLM can load GGUF, but on NVIDIA GPUs it's better to use official FP8 or NVFP4 weights, because they use the Tensor Cores of newer cards. How to choose between FP8, NVFP4 and AWQ is covered in model quantization on Blackwell GPUs.
How do you secure a model server?
One rule for every engine: the model server does not sit directly on the user network or on the internet.
- Ollama listens only on 127.0.0.1:11434 by default and has no authentication of its own. Changing
OLLAMA_HOSTto 0.0.0.0 with nothing in front opens the model to everyone on the network. - llama.cpp and vLLM have an API key option (
--api-key), but it's a shared key, with no users, roles or log. - Put a gateway in front of the model: user authentication, limits, a request log and one address for all applications.
The wider risks of LLM applications are covered in our article on the OWASP Top 10 for LLMs.
How we do it
On our servers with NVIDIA Blackwell GPUs we serve models with vLLM and SGLang. Applications see one OpenAI-compatible address, so swapping the engine or the model underneath needs no changes in the applications. That lets us pick, for each model, the engine that performs better on our traffic and switch it without downtime for users. We deploy models to the GPU fleet with our own tool, inferctl, which keeps Ansible and Terraform in one repository.
We use Ollama and llama.cpp where they have the edge: on laptops, for a quick look at a new model and for prototypes.
The most important lesson from practice: an engine upgrade is a production change, not a formality. We've seen a new engine version, a different tool-call parser or enabling speculative decoding break agents' tool calls while ordinary chat kept working fine. So before every version change we run a test suite on real agent tasks, and the previous container image stays on the server for a quick rollback.
Checklist for choosing
- How many concurrent users? One or a few people on their own hardware: Ollama or llama.cpp. A company API for a dozen or more people: vLLM or SGLang.
- What hardware? A server with NVIDIA or AMD GPUs: vLLM or SGLang. A laptop, Mac, CPU: Ollama or llama.cpp.
- How long are the prompts? RAG and agents with long context need good KV cache management and prefix caching.
- Do you need agents? Check the tool-call parser for your model in both engines and test it on your tools.
- Will there be LoRA adapters? Both engines support many adapters on one base model.
- vLLM or SGLang? Measure both on your traffic and your model before choosing.
- Who will maintain it? Pick an engine your team can upgrade and debug.
- What sits in front of the model? A gateway with authentication and logs, before the first user gets the address.
How all this compares with an API on cost is worked out in how much does an on-prem LLM cost. How we design the whole model-serving layer is described on our AI infrastructure page.
Sources
Engine documentation:
- vLLM: repository and feature list
- vLLM: OpenAI-Compatible Server
- vLLM: Automatic Prefix Caching
- vLLM: LoRA Adapters
- SGLang: repository and feature list
- SGLang: Hierarchical KV Caching (HiCache)
- SGLang: LoRA Serving
- SGLang: Quantization
- SGLang: Tool Parser
- SGLang: Anthropic-Compatible API
- Ollama: FAQ
- Ollama: OpenAI compatibility
- llama.cpp: README and server README
- Hugging Face: Text Generation Inference (maintenance mode)
Papers:
