This is the second part of the series. In Part 1 we work out weight and KV cache memory step by step and cover the cards from the RTX 5060 Ti to the RTX PRO 4000.
Quick answer
96 GB (RTX PRO 6000 Blackwell) is the threshold for company RAG and an agent backend: 30B models for dozens of users, ~70B for a few, 100-120B MoE. A single B300 (288 GB HBM3E, 8 TB/s) serves ~70B to many users or a 235B-class MoE in FP4. Frontier-class open models, from 0.7 to 2.8 trillion parameters, need an 8×B300 node (2.1 TB) or several nodes. Such a node can train a small model from scratch and fine-tune a large one, but training a frontier-level model from scratch is a job for clusters with thousands of GPUs.
What fits on which GPU?
The table reflects the state of things on 14 September 2026. "Realistically" means weights plus a KV cache for a sensible context, with about 10% headroom.
| Tier | VRAM | Memory bandwidth | What realistically fits | Typical use |
|---|---|---|---|---|
| RTX PRO 6000 Blackwell | 96 GB GDDR7 with ECC | 1792 GB/s (Server Edition: 1597 GB/s) | 24-32B in FP8 with lots of room, ~70B in FP8 or FP4, 100-120B MoE (gpt-oss-120b, Mistral Small 4 in NVFP4) | RAG and agents for a company, dozens of users on a 30B model, QLoRA of 70B models |
| B300 (single GPU) | 288 GB HBM3E | 8 TB/s | ~70B in BF16 or FP8, ~235B MoE in FP4, DeepSeek-V4-Flash | serving models to many users, full fine-tuning of ~8B models |
| 8×B300 (HGX/DGX B300) | 2.1 TB HBM3E total | 8 × 8 TB/s, NVLink 1.8 TB/s per GPU | open MoE models of 0.7-1.6T parameters in FP8 or FP4, 2.4-2.8T only in FP4 | frontier-class models in production, full fine-tuning up to ~70B, training small models from scratch |
At pimento we run models on NVIDIA Blackwell, but this article does not describe our setup. It is a general market guide based on NVIDIA specifications and model cards on Hugging Face.
The method in brief
The full calculation, with a worked config.json example, is in Part 1, in the section on working out memory. Here are just the rules we use below:
- Weights ≈ parameters × bytes per parameter: BF16 2 bytes, FP8 1 byte, FP4 about 0.6 bytes (from the real sizes of NVFP4 and MXFP4 files).
- KV cache per token = 2 × layers × KV heads × head dimension × bytes. For Llama 3.3 70B that is about 0.33 MB per token in BF16, or about 10.7 GB for 32k tokens. An FP8 cache takes half.
- Usable memory ≈ 90% of VRAM, and cache memory is usable memory minus weights.
- MoE: total parameters decide memory, active parameters decide speed.
- Speed: the tokens-per-second ceiling for one conversation ≈ memory bandwidth ÷ bytes of active weights. Llama 3.3 70B in FP8 (~71 GB) has a ceiling of ~25 tokens/s on an RTX PRO 6000 and ~113 tokens/s on a B300.
What are the specs of these GPUs?
All figures come from NVIDIA product pages, datasheets and whitepapers. We do not give prices: NVIDIA's pages do not state them, and prices change weekly.
| GPU | VRAM | Bandwidth | FP4 / FP8 | GPU-to-GPU link | Form factor and power |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell Workstation Edition | 96 GB GDDR7 with ECC, 512-bit | 1792 GB/s | yes / yes, 4000 AI TOPS | PCIe 5.0, NVIDIA does not list NVLink | dual slot, 600 W |
| RTX PRO 6000 Blackwell Max-Q | 96 GB GDDR7 with ECC, 512-bit | 1792 GB/s | yes / yes, 3511 AI TOPS | PCIe 5.0 x16, up to 4 cards per workstation | dual slot, 300 W |
| RTX PRO 6000 Blackwell Server Edition | 96 GB GDDR7 with ECC, 512-bit | 1597 GB/s | 4 PFLOPS FP4, 2 PFLOPS FP8 | PCIe 5.0, NVIDIA does not list NVLink | server, passive or liquid cooling, up to 600 W (configurable) |
| B300 (Blackwell Ultra) | 288 GB HBM3E | 8 TB/s | 15 PFLOPS FP4 (NVFP4, dense), 5 PFLOPS FP8 (dense) | NVLink 5: 1.8 TB/s per GPU; host: PCIe 6.0 x16 | SXM module on an HGX board, up to 1,400 W |
| HGX/DGX B300 (8 GPUs) | 2.1 TB HBM3E total | 8 × 8 TB/s (our calculation, NVIDIA does not give a total) | 108 PFLOPS FP4 dense, 72 PFLOPS FP8 with sparsity | NVLink 5 through switches, 14.4 TB/s total | DGX B300: 10U, about 14 kW |
A few notes on the table:
- "AI TOPS" on RTX PRO cards are FP4 with sparsity, a theoretical figure. The RTX Blackwell whitepaper confirms FP8 and FP4 support in the Tensor Cores of both consumer and professional cards.
- B300 memory. NVIDIA's technical blog gives 288 GB of HBM3E per GPU, while the HGX and DGX B300 pages give 2.1 TB for 8 GPUs, about 262 GB per GPU. NVIDIA does not explain the difference. For a single B300 we use 288 GB, and 2.1 TB for the node.
- B300 compute. The blog gives 15 PFLOPS FP4 (dense) per GPU, and the HGX page 108 PFLOPS FP4 (dense) for 8 GPUs, or 13.5 per GPU. NVIDIA does not explain this difference either.
- The B300 is not a card you put in a PC. It is an SXM module sold in HGX/DGX systems (8 GPUs) or in GB300 NVL72 racks. The "single B300" tier in this article means one GPU from such a system, for example rented in the cloud.
RTX PRO 6000 Blackwell 96 GB: why is this the threshold for a company?
96 GB with ECC and 1792 GB/s (Server Edition 1597 GB/s). Three editions: Workstation Edition (600 W), Max-Q (300 W, up to 4 cards per workstation) and Server Edition (passive, for servers, up to 8 cards in an RTX PRO Server). The card can also be split into isolated MIG instances: up to 4 × 24 GB, 2 × 48 GB or 1 × 96 GB, so you can run, for example, four independent 8B models. Usable memory is about 86 GB.
| Model class | Fits? | Precision | Realistic context and concurrency |
|---|---|---|---|
| dense 7-14B | yes, easily | BF16 or FP8 | 7-9B in BF16: ~475k tokens in a BF16 cache, dozens of conversations or 4 MIG instances. 14B in FP8: ~440k |
| dense 24-32B | yes | FP8 (32.8 GB) or BF16 (65.6 GB) | FP8: ~200k tokens in BF16, ~400k in FP8. For example 25 conversations of 16k each |
| dense ~70B (Llama 3.3 70B) | yes | FP8 (~71 GB) or FP4 (~42 GB) | FP8: ~48k tokens in BF16, ~96k in FP8, a few people. FP4: ~134k in BF16 |
| MoE ~30B-A3B | yes, easily | BF16 or FP8 | FP8: ~570k tokens in BF16. Dozens of users |
| MoE ~100-120B (gpt-oss-120b, Mistral Small 4, GLM-4.5-Air) | yes | gpt-oss-120b in MXFP4 (65.2 GB), Mistral Small 4 in NVFP4 (70.8 GB), GLM-4.5-Air in FP4 (~64 GB) | gpt-oss-120b: ~21 GB for cache at ~0.04 MB per token. GLM-4.5-Air: ~120k tokens. A few to a dozen people |
| MoE ~235B | not on one card | FP4 is ~141 GB | 2 cards in FP4, 4 cards in FP8 |
| frontier 0.7-1T+ | not on one card | NVFP4 is ~400-415 GB | 8 cards (768 GB) in an RTX PRO Server, over PCIe |
What you can actually do:
- RAG for a team and a company: yes. A 30B model in FP8 or a 30B-A3B MoE serves dozens of people with a 16k-token context.
- Agent backend for a company: yes, on 30B-120B models. gpt-oss-120b fits entirely on one card with a large context budget.
- LoRA: on a BF16 base up to ~32B. QLoRA of a 70B model: the QLoRA authors fine-tuned a 65B model on a single 48 GB GPU.
- Full fine-tuning: models of around 3-4B (96 GB ÷ 16 bytes per parameter is 6B before you count activations).
- Many users: dozens on models up to 30B.
- Several cards: 2 × 96 GB fits Qwen3-235B-A22B in FP4, 4 × Max-Q (384 GB) fits it in FP8 with ~110 GB for cache, and 8 cards in an RTX PRO Server (8 × 96 = 768 GB) fit ~700B models in NVFP4, such as DeepSeek-V3.1 (413 GB) or Mistral Large 3 (403 GB), with ~280 GB for cache. All of this, however, runs over PCIe, without NVLink.
B300: what does a single data-center GPU change?
288 GB of HBM3E and 8 TB/s, three times the memory and 4.5 times the bandwidth of the RTX PRO 6000. Plus NVLink 5 (1.8 TB/s per GPU) to the other GPUs in the node. Usable memory is about 259 GB.
| Model class | Fits? | Precision | Realistic context and concurrency |
|---|---|---|---|
| dense 7-14B | yes | BF16 | over 1.4 million tokens of cache. Hundreds of conversations |
| dense 24-32B | yes | BF16 or FP8 | FP8: ~860k tokens in BF16 |
| dense ~70B | yes | BF16 (~141 GB) or FP8 (~71 GB) | FP8: ~575k tokens in BF16, ~1.15 million in FP8. For example 70 conversations of 16k each. Generation ceiling ~113 tokens/s per conversation |
| MoE ~30B-A3B | yes | BF16 | over 2 million tokens |
| MoE ~100-120B | yes | FP8 or FP4 | GLM-4.5-Air in FP8 (~113 GB): ~810k tokens |
| MoE ~235B (Qwen3-235B-A22B, DeepSeek-V4-Flash) | yes | FP4 (~141 GB) for Qwen3-235B. FP8 (236 GB) leaves ~24 GB, too little for many people. DeepSeek-V4-Flash: official files 167 GB | Qwen3-235B in FP4: ~610k tokens in BF16 |
| frontier 0.7-1T+ | no | NVFP4 is ~400 GB and up | needs at least 2 GPUs, in practice a node |
What you can actually do:
- Serving many users: this is what the tier is for. High bandwidth gives fast responses, and the large memory holds the cache of hundreds of conversations.
- Agent backend for a company: models up to a 235B MoE on one GPU.
- LoRA: on a BF16 base up to ~70B (141 GB of weights plus activations).
- Full fine-tuning: ~8B models (16 bytes × 8.2B is about 131 GB plus activations), ~14B at the limit.
- Training from scratch: small experimental models, not production-class ones (details in the cluster section).
The 8×B300 cluster: what fits on an HGX or DGX B300 node?
An HGX B300 or DGX B300 node is 8 Blackwell Ultra GPUs, 2.1 TB of HBM3E, NVLink 5 with switches (1.8 TB/s per GPU, 14.4 TB/s total), 108 PFLOPS FP4 (dense) and about 14 kW in a 10U chassis (DGX). Usable memory is about 1.9 TB. This is the first tier on which frontier-class open models fit entirely:
| Model | Parameters (total, active) | Official files | On 8×B300 |
|---|---|---|---|
| DeepSeek-V3.1 | 671B, 37B | 689 GB (FP8) | yes, ~1.2 TB for cache |
| Mistral Large 3 | 675B, 41B | 682 GB (FP8), 403 GB (NVFP4) | yes. Mistral lists FP8 on a single B200 or H200 node |
| GLM-5 | 744B, 40B | 756 GB (FP8 version) | yes, ~1.1 TB for cache |
| Kimi K2.6 | 1T, 32B | 595 GB (native INT4) | yes, ~1.3 TB for cache |
| DeepSeek-V4-Pro-0813 | 1.6T, 49B | 893 GB | yes. DeepSeek shows it running on a single 4×GB300 node |
| Kimi K3 | 2.8T, 104B | 1.56 TB (MXFP4) | yes, ~330 GB for cache |
| Qwen3.8-2.4T-A95B | 2.4T, 95B | 2.5 TB (FP8 version) | not in FP8. It needs FP4 quantization (~1.4-1.5 TB by our estimate) or two nodes |
The DeepSeek-V4-Pro parameter counts come from the table in the DeepSeek-V4.1-Flash model card (1.6T "backbone" parameters, 49B active), because the V4-Pro-0813 card does not state them. File sizes are the sum of the safetensors files in the Hugging Face repositories.
The cache in these models is small thanks to MLA: for Kimi K2 and DeepSeek-V3 it is about 0.07 MB per token (from config.json: 61 layers × (512 + 64) × 2 bytes). With ~1.2 TB of free memory that means millions of cached tokens, or hundreds of parallel long-context agent conversations.
What you can actually do on 8×B300:
- Serving frontier models to many users: yes, this is the main use.
- LoRA and QLoRA of large MoE models: yes. OpenAI says gpt-oss-120b can be fine-tuned on a single H100 node (8 × 80 GB), and a B300 node has more than three times that memory.
- Full fine-tuning: dense models up to about 70B (16 bytes × 70B is ~1.1 TB of model state, sharded across 8 GPUs with FSDP or ZeRO-3, plus activations).
- Training from scratch: small models yes, frontier-level models no. The numbers are below.
How long would training from scratch take on one node?
Training cost is estimated with C ≈ 6 × N × D operations (N is parameters, D is training tokens) (Kaplan et al., 2020). HGX B300 lists 36 PFLOPS in BF16 with sparsity, which is 18 PFLOPS dense. We assume training uses 40% of that (our assumption, typical of well-optimized training, not a measurement), or ~7.2 PFLOPS.
| Model and data | Compute | Time on one 8×B300 (BF16, 40%) |
|---|---|---|
| 1B on 20 billion tokens (~20 tokens per parameter, as in Chinchilla) | 1.2 × 10²⁰ | ~4.6 hours |
| 7B on 140 billion tokens | 5.9 × 10²¹ | ~9.5 days |
| 8B on 15 trillion tokens (current practice) | 7.2 × 10²³ | ~3.2 years |
| 70B on 15 trillion tokens | 6.3 × 10²⁴ | ~28 years |
Training in FP8 roughly doubles the pace, but the order of magnitude stays. For comparison with model cards: Llama 3.3 70B (trained on over 15 trillion tokens) used 7.0 million H100 GPU hours, DeepSeek-V3 2.788 million H800 GPU hours, and Mistral Large 3 was trained on 3,000 H200 GPUs. The conclusion: one node can train a small model from scratch for research or a narrow task, but a model that competes with the frontier is a job for clusters of thousands of GPUs on a fast network.
Where does NVLink matter?
How splitting a model across several GPUs works (tensor parallelism and pipeline parallelism) is covered in Part 1. In short: PCIe 5.0 x16 gives 128 GB/s (NVIDIA, H100), while NVLink 5 on the B300 gives 1.8 TB/s per GPU, about 14 times more. The vLLM documentation recommends tensor parallelism when the model fits within one node, and pipeline parallelism for GPUs without NVLink (vLLM).
- RTX PRO 6000: NVIDIA does not list NVLink in its specification. Scaling goes over PCIe. The 8-card RTX PRO Server uses ConnectX-8 cards with a built-in PCIe 6.0 switch, which NVIDIA says doubles inter-GPU bandwidth compared with discrete PCIe switches, but it is still not NVLink.
- B300 in HGX/DGX: NVLink switches connect all 8 GPUs, so tensor parallelism across 8 GPUs is standard, and nodes connect over an 800 Gb/s network (ConnectX-8).
How much memory does fine-tuning need at these tiers?
The math is the same as in Part 1: QLoRA keeps the frozen base in 4 bits (Dettmers et al., 2023), LoRA in BF16 (Hu et al., 2021), and full fine-tuning with mixed-precision Adam needs 16 bytes per parameter, before activations (Rajbhandari et al., ZeRO).
| Tier | QLoRA | LoRA (BF16 base) | Full fine-tuning |
|---|---|---|---|
| RTX PRO 6000 96 GB | ~70B | up to ~32B | ~3-4B |
| B300 288 GB | large MoE up to ~235B | up to ~70B | ~8B, 14B at the limit |
| 8×B300 2.1 TB | frontier-class models | ~235B MoE in BF16 | dense up to ~70B |
These are estimates from the formulas, rounded down, because activations grow with sequence length and batch size. Gradient checkpointing, 8-bit optimizers and offloading optimizer state to RAM push these limits up at the cost of time.
How do you choose between an RTX PRO 6000, a B300 and a cluster?
Start with the task, not the card: what model quality you need, how many people will use it at once, and how long a context you process.
| Your situation | Sensible tier |
|---|---|
| RAG and agents for a company, dozens of users, data must stay in-house | RTX PRO 6000 Blackwell (1-4 cards) |
| A 235B or ~700B MoE on your own hardware without NVLink | 2-8 × RTX PRO 6000 with pipeline parallelism, with lower throughput than on B300 |
| A production service for hundreds of users, 70-235B models | B300 GPUs (owned or rented) |
| Frontier-class open models, full fine-tuning up to 70B | an 8×B300 node |
| Your own model, trained from scratch to compete with the frontier | a multi-node cluster, in practice thousands of GPUs |
If a personal assistant or RAG for a few people is enough, the 16 to 32 GB cards are covered in Part 1.
We show how we design AI on your own infrastructure, from choosing GPUs to serving models, on our AI infrastructure page. When it makes sense to move from an API to your own model at all, we cover in should you build or buy AI.
Sources
GPU specifications:
- NVIDIA: RTX Blackwell GPU Architecture (whitepaper)
- NVIDIA: RTX Blackwell PRO GPU Architecture (whitepaper)
- NVIDIA: RTX PRO 6000 Blackwell Workstation Edition
- NVIDIA: RTX PRO 6000 Blackwell Max-Q datasheet
- NVIDIA: RTX PRO 6000 Blackwell Server Edition
- NVIDIA: RTX PRO Server
- NVIDIA: Inside NVIDIA Blackwell Ultra
- NVIDIA: HGX
- NVIDIA: DGX B300
- NVIDIA: H100 (PCIe Gen5 bandwidth)
Model cards and configurations:
- Qwen/Qwen3-8B
- Qwen/Qwen3-14B
- Qwen/Qwen3-32B
- Qwen/Qwen3-30B-A3B-Instruct-2507
- Qwen/Qwen3-235B-A22B-Instruct-2507 and FP8 version
- Qwen/Qwen3.8-2.4T-A95B and FP8 version
- mistralai/Mistral-Small-4-119B-2603 and NVFP4 version
- mistralai/Mistral-Large-3-675B-Instruct-2512 and NVFP4 version
- meta-llama/Llama-3.3-70B-Instruct
- openai/gpt-oss-120b
- zai-org/GLM-4.5-Air and FP8 version
- zai-org/GLM-5 and FP8 version
- deepseek-ai/DeepSeek-V3
- deepseek-ai/DeepSeek-V3.1 and NVFP4 by NVIDIA
- deepseek-ai/DeepSeek-V4-Pro-0813
- deepseek-ai/DeepSeek-V4-Flash-0731
- deepseek-ai/DeepSeek-V4.1-Flash
- moonshotai/Kimi-K2-Instruct
- moonshotai/Kimi-K2.6
- moonshotai/Kimi-K3
Method, serving and training:
- vLLM: Parallelism and Scaling
- Dettmers et al.: QLoRA, Efficient Finetuning of Quantized LLMs (2023)
- Hu et al.: LoRA, Low-Rank Adaptation of Large Language Models (2021)
- Rajbhandari et al.: ZeRO, Memory Optimizations Toward Training Trillion Parameter Models (2019)
- Kaplan et al.: Scaling Laws for Neural Language Models (2020)
- Hoffmann et al.: Training Compute-Optimal Large Language Models (Chinchilla, 2022)
