Quick answer

Ollama and llama.cpp are tools for local work: a laptop, a home server, a single workstation, a prototype, one user or a few people. They are very good at it. vLLM and SGLang are production engines: you put them behind a company API that dozens of people and agents use at the same time. You won't see the difference with one question at a time. You'll see it when a dozen people ask at once and an agent sends long prompts full of documents. We serve models with vLLM and SGLang, and we decide which of them runs a given model by measuring.

As of 20 August 2026: vLLM 0.27.1, SGLang 0.5.17, Ollama 0.32.15, and llama.cpp released as daily builds.

Two worlds: local and production

Ollamallama.cpp (llama-server)vLLMSGLang
Built forlaptop, home, prototype, a few peoplesingle user, any hardwarecompany API, many users and agentscompany API, agents, long shared prefixes
HardwareCPU, GPU, MacCPU, NVIDIA, AMD and Intel GPUs, Vulkan, MacNVIDIA, AMD and Intel GPUs, CPU, others via pluginsNVIDIA and AMD GPUs, Intel Xeon CPUs, TPUs, Ascend NPUs
Weight formatGGUF (Ollama library)GGUFHugging Face safetensors (BF16, FP8, NVFP4, AWQ, GPTQ and more), also GGUFHugging Face safetensors, FP4, FP8, INT4, AWQ, GPTQ quantizations
Concurrent requests1 per model by defaultslots, continuous batching on by defaultcontinuous batching, PagedAttentioncontinuous batching, RadixAttention
Prompt reusenot described in the FAQprompt cache per slotautomatic prefix cachingprefix tree (RadixAttention), optional cache in RAM and on disk (HiCache)
APIown + OpenAI-compatibleOpenAI-compatibleOpenAI-compatible, Anthropic Messages, gRPCOpenAI-compatible, Anthropic Messages
LicenseMITMITApache 2.0Apache 2.0

Sources for the table: Ollama FAQ, llama.cpp server README, vLLM README, SGLang README, SGLang: HiCache, SGLang: Anthropic-Compatible API.

Why does the number of users change everything?

A language model generates its answer token by token. For every token the GPU has to read all the model's weights from memory. When one person asks, the card spends most of its time waiting for memory and its compute sits idle. When a dozen people ask at once, the engine can process their requests in one batch: the weights are read once and a dozen answers are computed.

The catch is the KV cache, the memory that holds each conversation's context. It grows with prompt length and with the number of concurrent requests. How to work it out is shown in our guide how much VRAM for a local LLM.

Engines differ in how they manage that cache:

  • Reserving up front. Each slot gets a fixed chunk of memory for the maximum context. Simple, but it wastes memory, because most conversations are shorter than the maximum. That's how local tools work, where one person is asking anyway.
  • PagedAttention (vLLM). The KV cache is split into small blocks allocated on demand, like memory pages in an operating system. The vLLM authors showed this delivers 2-4 times the throughput at the same latency compared with earlier systems (Kwon et al., 2023).
  • RadixAttention (SGLang). The KV cache of all requests is kept in a prefix tree (radix tree). A new request that starts the same way as an earlier one reuses the part already computed, and rarely used branches are evicted when memory runs short (Zheng et al., 2023). This suits agents well: they branch conversations and keep sending the same system prompt and the same tool definitions.
  • Continuous batching. A new request joins the running batch as soon as there's room, instead of waiting for every answer in the batch to finish.
  • Prefix caching. When many requests start the same way, the engine doesn't recompute the shared beginning. In vLLM this is automatic prefix caching (vLLM: Automatic Prefix Caching), in SGLang the tree described above. In RAG and agents this saves a lot of time to first token.

Ollama and llama.cpp: great locally

Ollama is the fastest way to run a model on your own computer: one command downloads the weights and starts a server. Under the hood it uses llama.cpp and the GGUF format. It has an OpenAI-compatible API, so an application written against Ollama can later move to a production engine without a rewrite (Ollama: OpenAI compatibility).

Why we don't put it behind a company API follows directly from the documentation (Ollama FAQ):

  • OLLAMA_NUM_PARALLEL defaults to 1, so one model handles one request at a time. Further requests go to a queue (up to 512 by default).
  • Raising parallelism increases the memory needed for context in proportion to the number of parallel requests and the context length.
  • Ollama loads models on demand and keeps several in memory (up to 3 per GPU by default). On a workstation that's convenient; on a production server it means unexpected reloads.

llama.cpp is the library and llama-server that Ollama, among others, is built on. It runs on CPUs, NVIDIA (CUDA), AMD (HIP) and Intel (SYCL) cards, via Vulkan and on Macs (Metal) (llama.cpp README), and GGUF quantizations let a model fit on hardware with little memory, at a cost in quality at low bit widths. The server can do more than people usually assume: slots with an automatically chosen count, continuous batching on by default, a shared KV cache buffer, a prompt cache, OpenAI-compatible endpoints, an API key and a router mode that loads several models from a directory (server README).

What to use them for:

  • a developer's laptop and a private assistant,
  • a home server or a single workstation,
  • learning, testing new models, prototypes,
  • Macs, CPUs, non-NVIDIA GPUs, edge devices,
  • models that need heavy quantization to fit at all.

What we don't trust them with: a company API used at the same time by a dozen or several dozen people and agents. That's where engines built around many concurrent conversations pull ahead.

vLLM and SGLang: production engines

When a model has to serve the whole company, meaning a dozen or several dozen people at once, agents and batch document processing, we pick one of two engines. Both are Apache 2.0 licensed and have a very similar feature set.

vLLM offers in one package (vLLM README):

  • PagedAttention, continuous batching, chunked prefill and prefix caching,
  • FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ and GGUF quantization,
  • speculative decoding,
  • tensor, pipeline, data and expert parallelism to split a model across several GPUs,
  • many LoRA adapters on one base model (vLLM: LoRA),
  • structured outputs plus tool-call and reasoning parsers for agents,
  • an OpenAI-compatible API, plus Anthropic Messages API and gRPC support,
  • support for NVIDIA, AMD and Intel GPUs, CPUs and other accelerators via plugins.

SGLang has a similar scope (SGLang README):

  • RadixAttention, i.e. prefix caching in a tree, continuous batching, paged attention and chunked prefill,
  • a hierarchical KV cache (HiCache): GPU, RAM and external storage, useful for long context and multi-turn conversations (SGLang: HiCache),
  • FP4, FP8, INT4, AWQ and GPTQ quantization (SGLang: Quantization),
  • speculative decoding and prefill-decode disaggregation onto separate machines,
  • tensor, pipeline, expert and data parallelism,
  • many LoRA adapters in one batch, based on S-LoRA and Punica techniques (SGLang: LoRA Serving),
  • structured outputs, tool and reasoning parsers (SGLang: Tool Parser),
  • an OpenAI-compatible API and an Anthropic-compatible /v1/messages endpoint, enabled by default (SGLang: Anthropic-Compatible API).

How do they differ in practice?

vLLMSGLang
KV cache managementPagedAttention, automatic prefix cachingRadixAttention (prefix tree), HiCache
Strong pointvery broad model and hardware support, frequent releasesworkloads with long shared prefixes: agents, multi-turn conversations
Configurationtool and reasoning parser chosen per model familytool and reasoning parser chosen per model family

Both projects move very fast and borrow ideas from each other, so feature differences blur from release to release. Performance differences depend on the model, the quantization, prompt lengths and the engine version. That's why we don't pick an engine from a ranking; we measure both on the traffic the model will serve.

The price of this power, for both: more configuration than Ollama, weights in Hugging Face format, frequent releases (a new version every few weeks), and new models often needing the latest engine version.

What about TGI?

Since December 2025, Hugging Face's Text Generation Inference has been in maintenance mode: it only accepts minor fixes and documentation. Hugging Face itself recommends vLLM, SGLang and local engines such as llama.cpp and MLX (TGI README). The last release, 3.3.7, dates from December 2025. It's not worth building new deployments on it.

Weight formats: GGUF or safetensors?

A practical difference that's easy to miss:

  • GGUF is the format of llama.cpp and Ollama. It has many quantization variants (e.g. Q4, Q5, Q8) and works well on CPUs and Macs.
  • Hugging Face safetensors is the format in which vendors publish models and official quantizations (FP8, NVFP4, AWQ). vLLM and SGLang work with it.

vLLM can load GGUF, but on NVIDIA GPUs it's better to use official FP8 or NVFP4 weights, because they use the Tensor Cores of newer cards. How to choose between FP8, NVFP4 and AWQ is covered in model quantization on Blackwell GPUs.

How do you secure a model server?

One rule for every engine: the model server does not sit directly on the user network or on the internet.

  • Ollama listens only on 127.0.0.1:11434 by default and has no authentication of its own. Changing OLLAMA_HOST to 0.0.0.0 with nothing in front opens the model to everyone on the network.
  • llama.cpp and vLLM have an API key option (--api-key), but it's a shared key, with no users, roles or log.
  • Put a gateway in front of the model: user authentication, limits, a request log and one address for all applications.

The wider risks of LLM applications are covered in our article on the OWASP Top 10 for LLMs.

How we do it

On our servers with NVIDIA Blackwell GPUs we serve models with vLLM and SGLang. Applications see one OpenAI-compatible address, so swapping the engine or the model underneath needs no changes in the applications. That lets us pick, for each model, the engine that performs better on our traffic and switch it without downtime for users. We deploy models to the GPU fleet with our own tool, inferctl, which keeps Ansible and Terraform in one repository.

We use Ollama and llama.cpp where they have the edge: on laptops, for a quick look at a new model and for prototypes.

The most important lesson from practice: an engine upgrade is a production change, not a formality. We've seen a new engine version, a different tool-call parser or enabling speculative decoding break agents' tool calls while ordinary chat kept working fine. So before every version change we run a test suite on real agent tasks, and the previous container image stays on the server for a quick rollback.

Checklist for choosing

  1. How many concurrent users? One or a few people on their own hardware: Ollama or llama.cpp. A company API for a dozen or more people: vLLM or SGLang.
  2. What hardware? A server with NVIDIA or AMD GPUs: vLLM or SGLang. A laptop, Mac, CPU: Ollama or llama.cpp.
  3. How long are the prompts? RAG and agents with long context need good KV cache management and prefix caching.
  4. Do you need agents? Check the tool-call parser for your model in both engines and test it on your tools.
  5. Will there be LoRA adapters? Both engines support many adapters on one base model.
  6. vLLM or SGLang? Measure both on your traffic and your model before choosing.
  7. Who will maintain it? Pick an engine your team can upgrade and debug.
  8. What sits in front of the model? A gateway with authentication and logs, before the first user gets the address.

How all this compares with an API on cost is worked out in how much does an on-prem LLM cost. How we design the whole model-serving layer is described on our AI infrastructure page.

Sources

Engine documentation:

Papers: