Quick answer
Quantization shrinks weights 2-4 times compared with BF16, so the same card fits a bigger model or more users. FP8 is the safe default: public studies find it effectively lossless. NVFP4 gives you almost another halving and native acceleration on Blackwell GPUs, but you have to check quality on your own tasks, especially for agents. AWQ and GPTQ are 4-bit weight-only quantizations, useful when memory is tight or your cards don't support FP8. Don't trust vendor claims or benchmarks on their own: measure on your own test set.
How do FP8, NVFP4, MXFP4 and AWQ differ?
| Format | What gets quantized | Bits per value | Hardware acceleration | Typical source of weights |
|---|---|---|---|---|
| BF16 | nothing (baseline) | 16 | every modern GPU | the vendor's original weights |
| FP8 (E4M3) | weights and activations (W8A8) | 8 | Tensor Cores from Ada and Hopper on, including Blackwell | official vendor releases, e.g. Qwen |
| NVFP4 | weights and activations (W4A4), often only some layers | ~4.5 | fifth-generation Tensor Cores (Blackwell) | NVIDIA Model Optimizer, some vendors, e.g. Mistral |
| MXFP4 | weights (e.g. MoE experts) | ~4.25 | Blackwell; fallback path on older cards | e.g. gpt-oss, released in MXFP4 from day one |
| AWQ, GPTQ | weights only (W4A16) | ~4 plus scales | no new cores needed, run from Turing on | the community, sometimes vendors |
| GGUF (Q4, Q5, Q8) | weights, various schemes | 4-8 | depends on the engine | the community, for llama.cpp and Ollama |
A few details that matter in practice:
- FP8 has 4 exponent bits and 3 mantissa bits (E4M3). The format is described in Micikevicius et al. (2022). Qwen publishes FP8 versions with fine-grained quantization in blocks of 128 and states that results are nearly identical to the original (Qwen3.8-27B-FP8).
- NVFP4 stores values as E2M1 (a range of roughly -6 to 6), in blocks of 16 values with one FP8 scale per block and a second FP32 scale per tensor. That gives about 4.5 bits per value and, according to NVIDIA, cuts memory by about 3.5 times versus FP16 and about 1.8 times versus FP8 (NVIDIA).
- MXFP4 is an open OCP standard with blocks of 32 values and a power-of-two scale (E8M0) (OCP Microscaling Formats). NVFP4's smaller blocks and fractional scale give a finer fit than MXFP4.
- AWQ picks out roughly 1% of the most important weight channels from activation statistics and scales them before quantization, reducing error without mixed precision (Lin et al., 2023, MLSys 2024 Best Paper). GPTQ quantizes weights layer by layer with error correction (Frantar et al., 2022).
How much memory a model takes in each format, and how much is left for the KV cache, is worked out in our guide how much VRAM for a local LLM.
Why "W8A8" and "W4A16" aren't the same thing
The label tells you what's quantized: W for weights, A for activations.
- W4A16 (AWQ, GPTQ) shrinks only the weights. When generating token by token, where the bottleneck is reading weights from memory, that helps. But the maths still runs in 16 bits, so with long prompts and many users at once, when the card is computing rather than waiting on memory, the gain shrinks.
- W8A8 (FP8) and W4A4 (NVFP4) shrink both weights and activations, so matrix multiplications run on lower-precision Tensor Cores. That also speeds up processing of long prompts, which is typical RAG and agent traffic.
A large study across the whole Llama 3.1 family (more than 500,000 evaluations) shows this well: FP8 W8A8 was effectively lossless, well-tuned INT8 lost 1-3%, and W4A16 did surprisingly well, on a par with 8-bit quantization. The authors recommend W4A16 for single, synchronous requests and W8A8 for continuous batching of many requests (Kurtic et al., ACL 2025).
What do public quality measurements say about NVFP4?
- NVIDIA reports that DeepSeek-R1-0528 quantized from FP8 to NVFP4 loses 1% or less on key tasks, for example MMLU-Pro 85% versus 84% and GPQA Diamond 81% versus 80% (NVIDIA).
- The nvidia/Qwen3.8-27B-NVFP4 model card shows results within about a percentage point of BF16: GPQA Diamond 88.92 versus 88.01, IFBench 80.07 versus 78.93, and AA-LCR even slightly higher (72.63 versus 73.38). An important detail: this is a mixed recipe. NVFP4 covers the MLP layers and the output head, while the attention layers stay in FP8.
These numbers are credible, but they measure what they measure: answers to test questions. They don't measure whether an agent calls the right tool with correct arguments, or whether the model holds a required format over hundreds of steps.
What we see in practice on Blackwell
On our NVIDIA Blackwell GPUs we serve models from the same families in both FP8 and NVFP4 variants and compare them on real agent and knowledge-base traffic. We don't publish our throughput measurements here. What follows are qualitative observations that changed our decisions:
- Plain text and maths unaffected, tool calls worse. In one test an NVFP4 variant wrote correct prose and got the maths right, but some of its tool calls were duplicated and in others the model invented arguments. The engine version we used had no native FP4 path for some of the layers and fell back to a weight-only kernel. Our conclusion: NVFP4 wasn't a drop-in replacement for FP8 behind an agent.
- Same engine, different output format. In an A/B comparison on the identical engine with identical prompts, the NVFP4 variant broke a strict output format the application relied on noticeably more often than the FP8 variant, which didn't break it at all. Since the engine and prompts were the same, the cause was the quantization. We switched back to FP8 with no changes on the client side, at the same address.
- Where "good text" is enough, NVFP4 works. A model from another family in NVFP4 served us steadily for a long time, captioning images during document indexing and answering in Polish.
- AWQ as a fallback. We used AWQ INT4 quantizations earlier for large dense models when memory mattered. Now that vendors publish official FP8 and NVFP4 releases, we reach for them less often.
That's why today we serve agents in FP8 by default, and choose NVFP4 for tasks where we've verified the difference doesn't matter.
Hardware and engine support: read the logs
"Blackwell" covers several chip variants. Data-center GPUs (such as the B200 and B300) and the desktop and workstation RTX cards have different compute architectures, and engines add kernels for them at different speeds. A public example: a vLLM issue from December 2025 describes how, on RTX Blackwell cards (SM120), an MXFP4 model didn't reach the native kernels and the engine fell back to the Marlin kernel, with a warning that this is weight-only quantization that may degrade performance (vLLM #31085).
What follows from that:
- After starting a model, read the engine logs. A warning about missing native FP4 support or about the Marlin kernel means you aren't getting what you chose a 4-bit format for.
- A fallback path can change not just speed but behaviour. That's exactly when our tool calls got worse.
- The hardware support table in the vLLM docs has no separate Blackwell column (vLLM: Quantization), so you have to verify support for your card and your engine version in practice.
- The engine matters too: vLLM and SGLang ship different kernels for the same formats. We compare engines in vLLM or SGLang.
Who made the quantization?
The same model in NVFP4 from three different authors is three different models. They differ in calibration data, method and which layers stayed in higher precision.
- The model vendor: Qwen publishes official FP8 releases, Mistral an official NVFP4 version of Mistral Small 4, and Google quantization-aware-trained (QAT) 4-bit Gemma 4 checkpoints that, according to Google, keep quality close to BF16 (Gemma 4 QAT).
- NVIDIA: publishes NVFP4 versions of popular models produced with Model Optimizer (NVIDIA Model Optimizer), usually with a results table in the card.
- The community: AWQ, GPTQ and GGUF releases. Some are excellent, but quality depends on the author. You can make your own with a tool such as LLM Compressor.
Order of preference: quantization-aware training or an official vendor release, then NVIDIA's with published results, and community releases last.
Which format when?
| Situation | Choice |
|---|---|
| Agents with tools, strict output format | FP8; NVFP4 only after an A/B test on your own tasks |
| RAG, summaries, image captions, chat | worth testing NVFP4, often good enough |
| The model doesn't fit in FP8 | NVFP4 (Blackwell) or AWQ/GPTQ (older cards) |
| Older GPUs without FP8 | AWQ or GPTQ |
| Single user, CPU, Mac | GGUF in llama.cpp or Ollama |
| Model released in 4 bits from the start (QAT, MXFP4) | the vendor's release, no further quantization |
A separate decision is the FP8 KV cache. It halves the memory needed for context, so you can serve more conversations or longer contexts (vLLM: Quantized KV Cache). Test that on long prompts too.
Checklist before changing format
- You have a baseline: the same model in FP8 or BF16 on the same engine.
- The test set covers your real tasks, not just general-knowledge questions.
- You measure tool calls separately: the right tool, correct arguments, no duplicates.
- You measure the required output format and language separately.
- You test long context, because quantization errors accumulate over long sessions.
- You check in the engine logs that it uses native kernels for the format.
- You know who made the quantization and what data it was calibrated on.
- You have a rollback plan: the previous weights and configuration stay ready for a quick switch.
How we build and run model serving on GPUs is described on our AI infrastructure page. How to choose the model itself before you decide on a format is covered in Qwen, Gemma, GLM, Mistral or Bielik.
Sources
Formats and methods:
- NVIDIA: Introducing NVFP4 for Efficient and Accurate Low-Precision Inference (2025)
- Open Compute Project: OCP Microscaling Formats (MX) v1.0
- Micikevicius et al.: FP8 Formats for Deep Learning (2022)
- Lin et al.: AWQ, Activation-aware Weight Quantization for LLM Compression and Acceleration (2023)
- Frantar et al.: GPTQ, Accurate Post-Training Quantization for Generative Pre-trained Transformers (2022)
- Kurtic et al.: "Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization (ACL 2025)
Model cards and tools:
- nvidia/Qwen3.8-27B-NVFP4
- Qwen/Qwen3.8-27B-FP8
- mistralai/Mistral-Small-4-119B-2603-NVFP4
- google/gemma-4-31B-it-qat-w4a16-ct
- NVIDIA Model Optimizer
- vLLM LLM Compressor
Engines:
