Quick answer

Quantization shrinks weights 2-4 times compared with BF16, so the same card fits a bigger model or more users. FP8 is the safe default: public studies find it effectively lossless. NVFP4 gives you almost another halving and native acceleration on Blackwell GPUs, but you have to check quality on your own tasks, especially for agents. AWQ and GPTQ are 4-bit weight-only quantizations, useful when memory is tight or your cards don't support FP8. Don't trust vendor claims or benchmarks on their own: measure on your own test set.

How do FP8, NVFP4, MXFP4 and AWQ differ?

FormatWhat gets quantizedBits per valueHardware accelerationTypical source of weights
BF16nothing (baseline)16every modern GPUthe vendor's original weights
FP8 (E4M3)weights and activations (W8A8)8Tensor Cores from Ada and Hopper on, including Blackwellofficial vendor releases, e.g. Qwen
NVFP4weights and activations (W4A4), often only some layers~4.5fifth-generation Tensor Cores (Blackwell)NVIDIA Model Optimizer, some vendors, e.g. Mistral
MXFP4weights (e.g. MoE experts)~4.25Blackwell; fallback path on older cardse.g. gpt-oss, released in MXFP4 from day one
AWQ, GPTQweights only (W4A16)~4 plus scalesno new cores needed, run from Turing onthe community, sometimes vendors
GGUF (Q4, Q5, Q8)weights, various schemes4-8depends on the enginethe community, for llama.cpp and Ollama

A few details that matter in practice:

  • FP8 has 4 exponent bits and 3 mantissa bits (E4M3). The format is described in Micikevicius et al. (2022). Qwen publishes FP8 versions with fine-grained quantization in blocks of 128 and states that results are nearly identical to the original (Qwen3.8-27B-FP8).
  • NVFP4 stores values as E2M1 (a range of roughly -6 to 6), in blocks of 16 values with one FP8 scale per block and a second FP32 scale per tensor. That gives about 4.5 bits per value and, according to NVIDIA, cuts memory by about 3.5 times versus FP16 and about 1.8 times versus FP8 (NVIDIA).
  • MXFP4 is an open OCP standard with blocks of 32 values and a power-of-two scale (E8M0) (OCP Microscaling Formats). NVFP4's smaller blocks and fractional scale give a finer fit than MXFP4.
  • AWQ picks out roughly 1% of the most important weight channels from activation statistics and scales them before quantization, reducing error without mixed precision (Lin et al., 2023, MLSys 2024 Best Paper). GPTQ quantizes weights layer by layer with error correction (Frantar et al., 2022).

How much memory a model takes in each format, and how much is left for the KV cache, is worked out in our guide how much VRAM for a local LLM.

Why "W8A8" and "W4A16" aren't the same thing

The label tells you what's quantized: W for weights, A for activations.

  • W4A16 (AWQ, GPTQ) shrinks only the weights. When generating token by token, where the bottleneck is reading weights from memory, that helps. But the maths still runs in 16 bits, so with long prompts and many users at once, when the card is computing rather than waiting on memory, the gain shrinks.
  • W8A8 (FP8) and W4A4 (NVFP4) shrink both weights and activations, so matrix multiplications run on lower-precision Tensor Cores. That also speeds up processing of long prompts, which is typical RAG and agent traffic.

A large study across the whole Llama 3.1 family (more than 500,000 evaluations) shows this well: FP8 W8A8 was effectively lossless, well-tuned INT8 lost 1-3%, and W4A16 did surprisingly well, on a par with 8-bit quantization. The authors recommend W4A16 for single, synchronous requests and W8A8 for continuous batching of many requests (Kurtic et al., ACL 2025).

What do public quality measurements say about NVFP4?

  • NVIDIA reports that DeepSeek-R1-0528 quantized from FP8 to NVFP4 loses 1% or less on key tasks, for example MMLU-Pro 85% versus 84% and GPQA Diamond 81% versus 80% (NVIDIA).
  • The nvidia/Qwen3.8-27B-NVFP4 model card shows results within about a percentage point of BF16: GPQA Diamond 88.92 versus 88.01, IFBench 80.07 versus 78.93, and AA-LCR even slightly higher (72.63 versus 73.38). An important detail: this is a mixed recipe. NVFP4 covers the MLP layers and the output head, while the attention layers stay in FP8.

These numbers are credible, but they measure what they measure: answers to test questions. They don't measure whether an agent calls the right tool with correct arguments, or whether the model holds a required format over hundreds of steps.

What we see in practice on Blackwell

On our NVIDIA Blackwell GPUs we serve models from the same families in both FP8 and NVFP4 variants and compare them on real agent and knowledge-base traffic. We don't publish our throughput measurements here. What follows are qualitative observations that changed our decisions:

  1. Plain text and maths unaffected, tool calls worse. In one test an NVFP4 variant wrote correct prose and got the maths right, but some of its tool calls were duplicated and in others the model invented arguments. The engine version we used had no native FP4 path for some of the layers and fell back to a weight-only kernel. Our conclusion: NVFP4 wasn't a drop-in replacement for FP8 behind an agent.
  2. Same engine, different output format. In an A/B comparison on the identical engine with identical prompts, the NVFP4 variant broke a strict output format the application relied on noticeably more often than the FP8 variant, which didn't break it at all. Since the engine and prompts were the same, the cause was the quantization. We switched back to FP8 with no changes on the client side, at the same address.
  3. Where "good text" is enough, NVFP4 works. A model from another family in NVFP4 served us steadily for a long time, captioning images during document indexing and answering in Polish.
  4. AWQ as a fallback. We used AWQ INT4 quantizations earlier for large dense models when memory mattered. Now that vendors publish official FP8 and NVFP4 releases, we reach for them less often.

That's why today we serve agents in FP8 by default, and choose NVFP4 for tasks where we've verified the difference doesn't matter.

Hardware and engine support: read the logs

"Blackwell" covers several chip variants. Data-center GPUs (such as the B200 and B300) and the desktop and workstation RTX cards have different compute architectures, and engines add kernels for them at different speeds. A public example: a vLLM issue from December 2025 describes how, on RTX Blackwell cards (SM120), an MXFP4 model didn't reach the native kernels and the engine fell back to the Marlin kernel, with a warning that this is weight-only quantization that may degrade performance (vLLM #31085).

What follows from that:

  • After starting a model, read the engine logs. A warning about missing native FP4 support or about the Marlin kernel means you aren't getting what you chose a 4-bit format for.
  • A fallback path can change not just speed but behaviour. That's exactly when our tool calls got worse.
  • The hardware support table in the vLLM docs has no separate Blackwell column (vLLM: Quantization), so you have to verify support for your card and your engine version in practice.
  • The engine matters too: vLLM and SGLang ship different kernels for the same formats. We compare engines in vLLM or SGLang.

Who made the quantization?

The same model in NVFP4 from three different authors is three different models. They differ in calibration data, method and which layers stayed in higher precision.

  • The model vendor: Qwen publishes official FP8 releases, Mistral an official NVFP4 version of Mistral Small 4, and Google quantization-aware-trained (QAT) 4-bit Gemma 4 checkpoints that, according to Google, keep quality close to BF16 (Gemma 4 QAT).
  • NVIDIA: publishes NVFP4 versions of popular models produced with Model Optimizer (NVIDIA Model Optimizer), usually with a results table in the card.
  • The community: AWQ, GPTQ and GGUF releases. Some are excellent, but quality depends on the author. You can make your own with a tool such as LLM Compressor.

Order of preference: quantization-aware training or an official vendor release, then NVIDIA's with published results, and community releases last.

Which format when?

SituationChoice
Agents with tools, strict output formatFP8; NVFP4 only after an A/B test on your own tasks
RAG, summaries, image captions, chatworth testing NVFP4, often good enough
The model doesn't fit in FP8NVFP4 (Blackwell) or AWQ/GPTQ (older cards)
Older GPUs without FP8AWQ or GPTQ
Single user, CPU, MacGGUF in llama.cpp or Ollama
Model released in 4 bits from the start (QAT, MXFP4)the vendor's release, no further quantization

A separate decision is the FP8 KV cache. It halves the memory needed for context, so you can serve more conversations or longer contexts (vLLM: Quantized KV Cache). Test that on long prompts too.

Checklist before changing format

  1. You have a baseline: the same model in FP8 or BF16 on the same engine.
  2. The test set covers your real tasks, not just general-knowledge questions.
  3. You measure tool calls separately: the right tool, correct arguments, no duplicates.
  4. You measure the required output format and language separately.
  5. You test long context, because quantization errors accumulate over long sessions.
  6. You check in the engine logs that it uses native kernels for the format.
  7. You know who made the quantization and what data it was calibrated on.
  8. You have a rollback plan: the previous weights and configuration stay ready for a quick switch.

How we build and run model serving on GPUs is described on our AI infrastructure page. How to choose the model itself before you decide on a format is covered in Qwen, Gemma, GLM, Mistral or Bielik.

Sources

Formats and methods:

Model cards and tools:

Engines: