Quick answer
GPU memory (VRAM) decides which AI model you can run locally, and memory bandwidth decides how fast it generates. A 16 GB card (RTX 5060 Ti) is enough for 7-9B models and a single-user assistant. 24 GB with ECC (RTX PRO 4000 Blackwell) handles RAG for a small team on an 8-14B model in FP8 and continuous operation in a server. 32 GB (RTX 5090) fits models up to about 32B in FP4 and small MoE models, with the same bandwidth as the RTX PRO 6000. ~70B models, MoE models over 100B and frontier-class open models need more memory: we cover them in Part 2.
What fits on which GPU?
The table reflects the state of things on 7 September 2026. "Realistically" means weights plus a KV cache for a sensible context, with about 10% headroom. The math is in the sections that follow.
| Tier | VRAM | Memory bandwidth | What realistically fits | Typical use |
|---|---|---|---|---|
| RTX 5060 Ti 16 GB | 16 GB GDDR7 | 448 GB/s | dense 7-9B in FP8 or FP4, 14B only in FP4, gpt-oss-20b | personal assistant, learning, prototypes, QLoRA of ~8B models |
| RTX PRO 4000 Blackwell | 24 GB GDDR7 with ECC | 672 GB/s | dense 7-14B in FP8, 30B-A3B MoE in FP4 with short context, 24-32B in FP4 for one person | RAG for a small team on an 8-14B model, dense low-power servers |
| RTX 5090 32 GB | 32 GB GDDR7 | 1792 GB/s | dense 7-14B in BF16 or FP8, 24-32B in FP4, 30B-A3B MoE in FP4 | strong personal assistant, RAG for a few people, LoRA of 7-8B models |
At pimento we run models on NVIDIA Blackwell, but this article does not describe our setup. It is a general market guide based on NVIDIA specifications and model cards on Hugging Face.
How do you work out how much memory a model needs?
GPU memory splits into three parts: model weights, the KV cache (the context memory) and overhead (CUDA context, activations, engine buffers). You can calculate or estimate each one before you buy hardware.
Weights: parameters times bytes per parameter
Weight memory ≈ number of parameters × bytes per parameter. The bytes depend on precision:
- BF16/FP16: 2 bytes. Most models are published this way.
- FP8: 1 byte. Some vendors publish weights directly in FP8 (for example Ministral 3, Mistral Small 4, DeepSeek-V3.1, GLM-5-FP8).
- FP4 (NVFP4, MXFP4): about 0.5-0.56 bytes. NVFP4 stores 4-bit values in blocks of 16, with one FP8 scale per block and one FP32 scale per tensor (NVIDIA). That is 4.5 bits, or about 0.56 bytes per parameter.
In practice FP4 files weigh a little more, because some layers (such as embeddings and attention) stay in higher precision. We checked the sizes of the official files on Hugging Face: gpt-oss-120b (117B parameters) weighs 65.2 GB, about 0.56 bytes per parameter, Mistral Small 4 in NVFP4 (119B) 70.8 GB (0.59), Mistral Large 3 in NVFP4 (675B) 403 GB (0.60), and NVIDIA's NVFP4 DeepSeek-V3.1 (671B) 413 GB (0.62). For estimates we therefore use 0.6 bytes per parameter in FP4. FP8 files stay close to 1 byte: Qwen3-235B-A22B-Instruct-2507-FP8 weighs 236 GB.
| Model | Parameters | BF16 | FP8 | FP4 (~0.6 B/param) |
|---|---|---|---|---|
| Qwen3-8B | 8.2B | 16.4 GB | 8.2 GB | ~4.9 GB |
| Qwen3-14B | 14.8B | 29.6 GB | 14.8 GB | ~8.9 GB |
| Qwen3-32B | 32.8B | 65.6 GB | 32.8 GB | ~19.7 GB |
| Llama 3.3 70B | 70B (70.6B per HF metadata) | ~141 GB | ~71 GB | ~42 GB |
| Qwen3-30B-A3B-Instruct-2507 | 30.5B, 3.3B active | 61 GB | 30.5 GB | ~18 GB |
| gpt-oss-120b | 117B, 5.1B active | not published | not published | 65.2 GB (official MXFP4) |
| Qwen3-235B-A22B-Instruct-2507 | 235B, 22B active | 470 GB | 236 GB (official FP8) | ~141 GB |
| DeepSeek-V3.1 | 671B, 37B active | ~1.34 TB | 689 GB (official) | 413 GB (NVFP4 by NVIDIA) |
FP8 is usually a safe choice for quality. FP4 has native support in the Tensor Cores of every Blackwell card in this series (NVIDIA RTX Blackwell, NVIDIA Blackwell Ultra), and more and more vendors train with 4 bits in mind from the start: gpt-oss shipped in MXFP4, Kimi K2.6 in native INT4, and Kimi K3 in MXFP4 with quantization-aware training. For models quantized after the fact, check FP4 quality on your own test set.
The KV cache: a worked example from Qwen3-32B's config.json
While generating, the model keeps the keys and values (K and V) from every layer for each context token. That is the KV cache. The formula:
KV per token = 2 (K and V) × number of layers × number of KV heads × head dimension × bytes per value.
Take Qwen3-32B's config.json: num_hidden_layers: 64, num_attention_heads: 64, num_key_value_heads: 8, head_dim: 128. The model uses GQA (grouped-query attention): 64 query heads share 8 KV heads. In BF16:
- 2 × 64 × 8 × 128 × 2 bytes = 262,144 bytes, or 256 KiB (about 0.26 MB) per token.
- A 32,768-token context: about 8.6 GB. A 128k-token context: about 34 GB.
- Without GQA (64 KV heads instead of 8) it would be 2 MiB per token, about 69 GB for 32k tokens. GQA cuts the cache eightfold.
- An FP8 cache takes half the space, and engines such as vLLM support it directly (vLLM).
The same formula (values from config.json, our calculation, BF16) gives for other models:
| Model | Layers, KV heads, head dimension | KV per token | 32k tokens |
|---|---|---|---|
| Qwen3-8B | 36, 8, 128 | ~0.15 MB | ~4.8 GB |
| Qwen3-14B | 40, 8, 128 | ~0.16 MB | ~5.4 GB |
| Llama 3.3 70B | 80, 8, 128 | ~0.33 MB | ~10.7 GB |
| Qwen3-30B-A3B | 48, 4, 128 | ~0.10 MB | ~3.2 GB |
| Qwen3-235B-A22B | 94, 4, 128 | ~0.19 MB | ~6.3 GB |
| gpt-oss-120b | 18 full-attention layers, 8, 64 | ~0.04 MB | ~1.2 GB |
Newer architectures shrink the cache further. gpt-oss interleaves full attention with a 128-token sliding window, so the cache grows in only half the layers. Gemma 4 uses a 1024-token window in 50 of its 60 layers, Qwen3.8-27B has full attention in 16 of 64 layers, and DeepSeek, Kimi K2 and GLM-5 compress the cache into a shared latent vector (MLA): for Kimi K2 that is about 0.07 MB per token. That is why you always calculate from the specific model's config.json.
Rule of thumb: context and number of users
- Usable memory ≈ 90% of VRAM. The rest goes to the CUDA context, activations and buffers. vLLM takes 92% of GPU memory by default (vLLM). Our 10% is an approximation, not a measurement.
- Cache memory = usable memory minus weights.
- Tokens in cache = cache memory ÷ KV per token.
- Concurrency ≈ tokens ÷ average context per conversation.
Example: Qwen3-32B in FP4 on an RTX 5090. 32 GB × 0.9 = 28.8 GB. Minus ~19.7 GB of weights leaves ~9 GB, or ~35k tokens in BF16 or ~70k in FP8. That is one person with a 32k context or four with a 16k context using an FP8 cache. An agent that loads a large repository or long documents will use the whole budget on its own.
MoE: total parameters decide memory, active parameters decide speed
An MoE (mixture of experts) model picks a few experts out of many for each token. Qwen3-30B-A3B has 128 experts, 8 of which work on each token. The router can pick any expert, so all the weights must be in VRAM: 30.5B parameters take as much memory as a dense 30B model. But only about 3.3B parameters are read per token, so the model generates at the pace of a ~3B model. That is why MoE is the best way to get high quality at high speed, provided you have the memory for all of it.
You can move some experts to system RAM (llama.cpp and similar tools do this). The model then runs, but each token reads part of the weights over PCIe from system memory, so generation slows several times or more. It is fine for experiments and poor for serving many people.
Why does memory bandwidth decide speed?
To generate each new token, the GPU has to read all the active weights from memory. There is little computation involved, so the decode phase is limited by memory bandwidth, not compute (NVIDIA). That gives a simple ceiling for a single conversation:
Maximum tokens per second ≈ memory bandwidth ÷ bytes of active weights.
- Qwen3-32B in FP4 (~19.7 GB) on an RTX 5090 (1792 GB/s): a ceiling of ~90 tokens/s.
- The same model on an RTX 5060 Ti (448 GB/s), if it fit: ~23 tokens/s.
- Llama 3.3 70B in FP8 (~71 GB) on an RTX PRO 6000 (1792 GB/s): ~25 tokens/s. On a B300 (8 TB/s): ~113 tokens/s.
These are upper bounds. Real results are lower, because the KV cache also has to be read and the engine adds overhead. Two conclusions still hold. First, the "AI TOPS" on the spec sheet (for example 3352 for the RTX 5090, counted in FP4 with sparsity) describe compute, which matters for prompt processing and many parallel conversations, not response speed for a single person. Second, when you serve many users at once, one read of the weights serves a whole batch of requests, so total throughput grows with the number of conversations. The condition: there has to be memory left for their KV cache.
What are the specs of these GPUs?
All figures come from NVIDIA product pages, datasheets and whitepapers. We do not give prices: NVIDIA's pages do not state them, and retail prices change weekly.
| GPU | VRAM | Bandwidth | FP4 / FP8 | GPU-to-GPU link | Form factor and power |
|---|---|---|---|---|---|
| RTX 5060 Ti 16 GB | 16 GB GDDR7, 128-bit | 448 GB/s | yes / yes (5th-gen Tensor Cores), 759 AI TOPS | PCIe 5.0 only, NVLink: no | desktop card, 180 W, 600 W PSU minimum |
| RTX PRO 4000 Blackwell | 24 GB GDDR7 with ECC | 672 GB/s | yes / yes, 1290 FP4 TOPS with sparsity | PCIe 5.0, NVIDIA does not list NVLink | single slot, 145 W |
| RTX 5090 | 32 GB GDDR7, 512-bit | 1792 GB/s | yes / yes, 3352 AI TOPS | PCIe 5.0 only, NVLink: no | desktop card, 575 W, 1000 W PSU minimum |
"AI TOPS" on RTX cards are FP4 with sparsity, a theoretical figure. On GeForce cards NVIDIA lists FP4 as a Tensor Core feature, and the RTX Blackwell whitepaper confirms FP8 and FP4 support in the Tensor Cores of both consumer and professional cards.
RTX 5060 Ti 16 GB: what is the cheapest tier good for?
16 GB and 448 GB/s is the tier for a personal assistant and learning. Usable memory is about 14.4 GB.
| Model class | Fits? | Precision | Realistic context and concurrency |
|---|---|---|---|
| dense 7-9B (Qwen3-8B, Ministral 3 8B) | yes | FP8 (8.2 GB) or FP4 (~4.9 GB) | FP8: ~42k tokens in a BF16 cache, ~84k in FP8. One or two people with 16-32k context |
| dense 14B (Qwen3-14B, Ministral 3 14B) | yes, tight | FP4 only (~8.9 GB) | ~34k tokens in BF16, ~67k in FP8. One person |
| dense 24-32B | no | FP4 is ~15-20 GB of weights alone | only with layers offloaded to RAM, a few tokens per second |
| MoE ~30B-A3B | partly | gpt-oss-20b (21B, 3.6B active) runs within 16 GB according to OpenAI; Qwen3-30B-A3B in FP4 (~18 GB) does not | gpt-oss-20b: short context, one person; Qwen3-30B-A3B only with experts in RAM |
| dense ~70B, MoE ~100-120B and ~235B, frontier 0.7-1T+ | no |
What you can actually do:
- Personal assistant: yes. An 8B model in FP8 has a generation ceiling of ~55 tokens/s, ~90 in FP4, which is faster than you can read.
- RAG for a team: only as a prototype for 1-2 people. Documents pushed into the context quickly use up 40-80k tokens of cache.
- Agent backend for a company: no. Agents need long context and parallel calls.
- LoRA of small models: QLoRA of a ~8B model yes (the 4-bit base takes ~5 GB). LoRA on a BF16 base no, because the 8B weights alone are 16.4 GB.
- Many users: no.
- Two cards: there is no NVLink, so you connect them over PCIe. 2 × 16 GB is still less than one 32 GB card, because part of each card's memory goes to overhead.
RTX PRO 4000 Blackwell 24 GB: who is an entry professional card for?
24 GB with ECC, 672 GB/s, 145 W, single slot. Bandwidth is lower than on the RTX 5090, but the card is slim, cool and has error-correcting memory, so it suits servers and workstations running around the clock. Usable memory is about 21.6 GB.
| Model class | Fits? | Precision | Realistic context and concurrency |
|---|---|---|---|
| dense 7-9B | yes | FP8 (8.2 GB), BF16 tight | FP8: ~91k tokens in BF16, ~182k in FP8. A few people |
| dense 14B | yes | FP8 (14.8 GB) or FP4 | FP8: ~42k tokens, FP4: ~78k (BF16 cache). 2-4 people |
| dense 24-32B | barely | FP4 (~19.7 GB for Qwen3-32B) | ~7k tokens in BF16, ~15k in FP8. One person, short context. Smaller 24B models (such as Mistral Small 3.2) leave more room |
| MoE ~30B-A3B | yes, tight | FP4 (~18 GB) | ~34k tokens in BF16, ~67k in FP8. gpt-oss-20b (13.8 GB) fits with room to spare |
| dense ~70B, MoE ~100-120B and ~235B, frontier 0.7-1T+ | no |
What you can actually do:
- Personal assistant: yes, with an 8-14B model.
- RAG for a team: yes, for a small team on an 8-14B model in FP8. The generation ceiling for 14B in FP8 is ~45 tokens/s per conversation.
- Agent backend: for simple agents on an 8-14B model.
- LoRA: QLoRA up to ~14B comfortably, ~24B tight. LoRA on a BF16 base for ~8B with short sequences.
- Many users: a handful on an 8B model.
- Several cards: low power and a single slot make it easier to fit several cards in one chassis, but only PCIe connects them. How many cards a given workstation takes depends on the system vendor.
RTX 5090 32 GB: how far can you push the strongest consumer card?
32 GB and 1792 GB/s: the same bandwidth as the RTX PRO 6000, but a third of the memory. 575 W and a 1000 W PSU minimum. Usable memory is about 28.8 GB.
| Model class | Fits? | Precision | Realistic context and concurrency |
|---|---|---|---|
| dense 7-9B | yes | BF16 (16.4 GB) or FP8 | FP8: ~140k tokens in BF16, ~280k in FP8. Several to a dozen conversations |
| dense 14B | yes | FP8 (14.8 GB) | ~85k tokens in BF16, ~170k in FP8 |
| dense 24-32B (Qwen3-32B, Gemma 4 31B, Qwen3.8-27B) | yes | FP4 (~18-20 GB) | ~35k tokens in BF16, ~70k in FP8. 1-4 people |
| dense ~70B | not on one card | FP4 is ~42 GB | on two cards (64 GB) in FP4, over PCIe |
| MoE ~30B-A3B (Qwen3-30B-A3B, Qwen3.6-35B-A3B, Gemma 4 26B A4B) | yes | FP4 (~18 GB for Qwen3-30B-A3B) | ~107k tokens in BF16, ~214k in FP8, because this model's cache is small (~0.1 MB per token) |
| MoE ~100-120B | no | gpt-oss-120b is 65.2 GB | only with experts in RAM |
| MoE ~235B, frontier 0.7-1T+ | no |
What you can actually do:
- Personal assistant: yes, with the best 27-32B models in FP4. Generation ceiling ~90 tokens/s.
- RAG for a team: for a few people at once. Best with a 30B-A3B MoE, which has a small cache and generates at the pace of a 3B model.
- Agent backend: as a trial or for a small team. A consumer card has no ECC and no rating for continuous data-center operation.
- LoRA: on a BF16 base for 7-8B, QLoRA up to ~32B with short sequences. OpenAI says gpt-oss-20b can be fine-tuned on consumer hardware.
- Many users: a dozen or so on an 8B model, a few on 32B.
- Two RTX 5090s: 64 GB lets you run 70B in FP4 with ~15 GB for cache. NVIDIA's specification states NVLink is not supported, so the cards talk over PCIe (details below).
Several cards: why is PCIe not NVLink?
When a model does not fit on one GPU, you split it across several. There are two basic ways:
- Tensor parallelism (TP): each layer is split across GPUs, which exchange results after every layer. That is dozens of synchronizations per token, so TP needs a very fast link.
- Pipeline parallelism (PP): each GPU gets a consecutive set of layers, and results move on only at the boundaries. There is less communication, but a single request does not get faster.
The gap between links is large. PCIe 5.0 x16 gives 128 GB/s (NVIDIA, H100), while NVLink 5 on the B300 gives 1.8 TB/s per GPU, about 14 times more. The vLLM documentation recommends TP when the model fits within one node, and for GPUs without NVLink (it gives the L40S as an example) recommends pipeline parallelism instead of tensor parallelism (vLLM).
What this means for the cards in this part:
- RTX 5060 Ti and RTX 5090: NVIDIA's specification says "NVLink: No". Two cards work over PCIe. That is enough to fit a larger model, but it does not give a linear speedup.
- RTX PRO 4000: NVIDIA does not list NVLink in its specification. Scaling goes over PCIe.
NVLink only appears in data-center GPUs such as the B300 in HGX/DGX systems. We cover them in Part 2.
How much memory does fine-tuning need?
Three methods, three different sums:
- QLoRA: the base is frozen in 4 bits, and you train only small adapters. The authors fine-tuned a 65B model this way on a single 48 GB GPU (Dettmers et al., 2023).
- LoRA: the base is frozen in BF16 (2 bytes per parameter) plus adapters. LoRA cuts the number of trainable parameters by up to 10,000 times and the GPU memory requirement by about 3 times compared with full fine-tuning (Hu et al., 2021).
- Full fine-tuning: with mixed-precision Adam that is 16 bytes per parameter (16-bit weights and gradients plus 12 bytes of optimizer state), before activations (Rajbhandari et al., ZeRO).
| Tier | QLoRA | LoRA (BF16 base) | Full fine-tuning |
|---|---|---|---|
| RTX 5060 Ti 16 GB | up to ~8B | no | no |
| RTX PRO 4000 24 GB | up to ~14B, ~24B tight | ~7-8B, short sequences | around 1B |
| RTX 5090 32 GB | up to ~32B, short sequences | ~7-8B | around 1B |
These are estimates from the formulas, rounded down, because activations grow with sequence length and batch size. Gradient checkpointing, 8-bit optimizers and offloading optimizer state to RAM push these limits up at the cost of time.
Training from scratch is not a job for this tier: we work out how long it would take even on an 8×B300 node in Part 2.
How do you choose a card from 16 to 32 GB?
Start with the task, not the card. Answer three questions: what model quality you need, how many people will use it at once, and how long a context you process.
| Your situation | Sensible tier |
|---|---|
| You are learning, prototyping, or need a private assistant | RTX 5060 Ti 16 GB with an 8B model, gpt-oss-20b |
| A strong personal assistant, coding with a 27-32B model | RTX 5090 32 GB |
| RAG for a small team, continuous operation, a low-power server | RTX PRO 4000 Blackwell (one or several cards) |
| RAG and agents for a company, dozens of users, 70B models and up | RTX PRO 6000, B300 or 8×B300, covered in Part 2 |
Three practical rules:
- Calculate from config.json, not from the model name. Two "30B" models can have KV caches several times apart, depending on GQA, MLA and attention windows.
- Leave memory for context. A model that "fits" on a card but leaves 2 GB for cache will serve one short conversation.
- For speed, look at memory bandwidth and active parameters. For capacity, look at VRAM and total parameters.
We show how we design AI on your own infrastructure, from choosing GPUs to serving models, on our AI infrastructure page. When it makes sense to move from an API to your own model at all, we cover in should you build or buy AI.
Part 2: RTX PRO 6000, B300 and an 8×B300 cluster → What fits in 96 GB, 288 GB and 2.1 TB, where NVLink comes in, and how long training from scratch would take on a single node.
Sources
GPU specifications:
- NVIDIA: GeForce RTX 5060 Family
- NVIDIA: GeForce RTX 5090
- NVIDIA: GeForce graphics cards, compare specs
- NVIDIA: RTX Blackwell GPU Architecture (whitepaper)
- NVIDIA: RTX Blackwell PRO GPU Architecture (whitepaper)
- NVIDIA: RTX PRO 4000 Blackwell
- NVIDIA: RTX PRO 6000 Blackwell Workstation Edition
- NVIDIA: Inside NVIDIA Blackwell Ultra
- NVIDIA: H100 (PCIe Gen5 bandwidth)
Model cards and configurations:
- Qwen/Qwen3-8B
- Qwen/Qwen3-14B
- Qwen/Qwen3-32B and config.json
- Qwen/Qwen3.8-27B
- Qwen/Qwen3-30B-A3B-Instruct-2507
- Qwen/Qwen3.6-35B-A3B
- Qwen/Qwen3-235B-A22B-Instruct-2507 and FP8 version
- mistralai/Ministral-3-8B-Instruct-2512
- mistralai/Ministral-3-14B-Instruct-2512
- mistralai/Mistral-Small-3.2-24B-Instruct-2506
- mistralai/Mistral-Small-4-119B-2603 and NVFP4 version
- mistralai/Mistral-Large-3-675B-Instruct-2512 and NVFP4 version
- google/gemma-4-31B-it
- google/gemma-4-26B-A4B-it
- meta-llama/Llama-3.3-70B-Instruct
- openai/gpt-oss-20b
- openai/gpt-oss-120b
- zai-org/GLM-5 and FP8 version
- deepseek-ai/DeepSeek-V3.1 and NVFP4 by NVIDIA
- moonshotai/Kimi-K2-Instruct
- moonshotai/Kimi-K2.6
- moonshotai/Kimi-K3
Method, serving and training:
- NVIDIA: Mastering LLM Techniques, Inference Optimization
- vLLM: Parallelism and Scaling
- vLLM: Quantized KV Cache
- vLLM: Engine Arguments
- Dettmers et al.: QLoRA, Efficient Finetuning of Quantized LLMs (2023)
- Hu et al.: LoRA, Low-Rank Adaptation of Large Language Models (2021)
- Rajbhandari et al.: ZeRO, Memory Optimizations Toward Training Trillion Parameter Models (2019)
