This is the second part of the series. In Part 1 we work out weight and KV cache memory step by step and cover the cards from the RTX 5060 Ti to the RTX PRO 4000.

Quick answer

96 GB (RTX PRO 6000 Blackwell) is the threshold for company RAG and an agent backend: 30B models for dozens of users, ~70B for a few, 100-120B MoE. A single B300 (288 GB HBM3E, 8 TB/s) serves ~70B to many users or a 235B-class MoE in FP4. Frontier-class open models, from 0.7 to 2.8 trillion parameters, need an 8×B300 node (2.1 TB) or several nodes. Such a node can train a small model from scratch and fine-tune a large one, but training a frontier-level model from scratch is a job for clusters with thousands of GPUs.

What fits on which GPU?

The table reflects the state of things on 14 September 2026. "Realistically" means weights plus a KV cache for a sensible context, with about 10% headroom.

TierVRAMMemory bandwidthWhat realistically fitsTypical use
RTX PRO 6000 Blackwell96 GB GDDR7 with ECC1792 GB/s (Server Edition: 1597 GB/s)24-32B in FP8 with lots of room, ~70B in FP8 or FP4, 100-120B MoE (gpt-oss-120b, Mistral Small 4 in NVFP4)RAG and agents for a company, dozens of users on a 30B model, QLoRA of 70B models
B300 (single GPU)288 GB HBM3E8 TB/s~70B in BF16 or FP8, ~235B MoE in FP4, DeepSeek-V4-Flashserving models to many users, full fine-tuning of ~8B models
8×B300 (HGX/DGX B300)2.1 TB HBM3E total8 × 8 TB/s, NVLink 1.8 TB/s per GPUopen MoE models of 0.7-1.6T parameters in FP8 or FP4, 2.4-2.8T only in FP4frontier-class models in production, full fine-tuning up to ~70B, training small models from scratch

At pimento we run models on NVIDIA Blackwell, but this article does not describe our setup. It is a general market guide based on NVIDIA specifications and model cards on Hugging Face.

The method in brief

The full calculation, with a worked config.json example, is in Part 1, in the section on working out memory. Here are just the rules we use below:

  • Weights ≈ parameters × bytes per parameter: BF16 2 bytes, FP8 1 byte, FP4 about 0.6 bytes (from the real sizes of NVFP4 and MXFP4 files).
  • KV cache per token = 2 × layers × KV heads × head dimension × bytes. For Llama 3.3 70B that is about 0.33 MB per token in BF16, or about 10.7 GB for 32k tokens. An FP8 cache takes half.
  • Usable memory ≈ 90% of VRAM, and cache memory is usable memory minus weights.
  • MoE: total parameters decide memory, active parameters decide speed.
  • Speed: the tokens-per-second ceiling for one conversation ≈ memory bandwidth ÷ bytes of active weights. Llama 3.3 70B in FP8 (~71 GB) has a ceiling of ~25 tokens/s on an RTX PRO 6000 and ~113 tokens/s on a B300.

What are the specs of these GPUs?

All figures come from NVIDIA product pages, datasheets and whitepapers. We do not give prices: NVIDIA's pages do not state them, and prices change weekly.

GPUVRAMBandwidthFP4 / FP8GPU-to-GPU linkForm factor and power
RTX PRO 6000 Blackwell Workstation Edition96 GB GDDR7 with ECC, 512-bit1792 GB/syes / yes, 4000 AI TOPSPCIe 5.0, NVIDIA does not list NVLinkdual slot, 600 W
RTX PRO 6000 Blackwell Max-Q96 GB GDDR7 with ECC, 512-bit1792 GB/syes / yes, 3511 AI TOPSPCIe 5.0 x16, up to 4 cards per workstationdual slot, 300 W
RTX PRO 6000 Blackwell Server Edition96 GB GDDR7 with ECC, 512-bit1597 GB/s4 PFLOPS FP4, 2 PFLOPS FP8PCIe 5.0, NVIDIA does not list NVLinkserver, passive or liquid cooling, up to 600 W (configurable)
B300 (Blackwell Ultra)288 GB HBM3E8 TB/s15 PFLOPS FP4 (NVFP4, dense), 5 PFLOPS FP8 (dense)NVLink 5: 1.8 TB/s per GPU; host: PCIe 6.0 x16SXM module on an HGX board, up to 1,400 W
HGX/DGX B300 (8 GPUs)2.1 TB HBM3E total8 × 8 TB/s (our calculation, NVIDIA does not give a total)108 PFLOPS FP4 dense, 72 PFLOPS FP8 with sparsityNVLink 5 through switches, 14.4 TB/s totalDGX B300: 10U, about 14 kW

A few notes on the table:

  • "AI TOPS" on RTX PRO cards are FP4 with sparsity, a theoretical figure. The RTX Blackwell whitepaper confirms FP8 and FP4 support in the Tensor Cores of both consumer and professional cards.
  • B300 memory. NVIDIA's technical blog gives 288 GB of HBM3E per GPU, while the HGX and DGX B300 pages give 2.1 TB for 8 GPUs, about 262 GB per GPU. NVIDIA does not explain the difference. For a single B300 we use 288 GB, and 2.1 TB for the node.
  • B300 compute. The blog gives 15 PFLOPS FP4 (dense) per GPU, and the HGX page 108 PFLOPS FP4 (dense) for 8 GPUs, or 13.5 per GPU. NVIDIA does not explain this difference either.
  • The B300 is not a card you put in a PC. It is an SXM module sold in HGX/DGX systems (8 GPUs) or in GB300 NVL72 racks. The "single B300" tier in this article means one GPU from such a system, for example rented in the cloud.

RTX PRO 6000 Blackwell 96 GB: why is this the threshold for a company?

96 GB with ECC and 1792 GB/s (Server Edition 1597 GB/s). Three editions: Workstation Edition (600 W), Max-Q (300 W, up to 4 cards per workstation) and Server Edition (passive, for servers, up to 8 cards in an RTX PRO Server). The card can also be split into isolated MIG instances: up to 4 × 24 GB, 2 × 48 GB or 1 × 96 GB, so you can run, for example, four independent 8B models. Usable memory is about 86 GB.

Model classFits?PrecisionRealistic context and concurrency
dense 7-14Byes, easilyBF16 or FP87-9B in BF16: ~475k tokens in a BF16 cache, dozens of conversations or 4 MIG instances. 14B in FP8: ~440k
dense 24-32ByesFP8 (32.8 GB) or BF16 (65.6 GB)FP8: ~200k tokens in BF16, ~400k in FP8. For example 25 conversations of 16k each
dense ~70B (Llama 3.3 70B)yesFP8 (~71 GB) or FP4 (~42 GB)FP8: ~48k tokens in BF16, ~96k in FP8, a few people. FP4: ~134k in BF16
MoE ~30B-A3Byes, easilyBF16 or FP8FP8: ~570k tokens in BF16. Dozens of users
MoE ~100-120B (gpt-oss-120b, Mistral Small 4, GLM-4.5-Air)yesgpt-oss-120b in MXFP4 (65.2 GB), Mistral Small 4 in NVFP4 (70.8 GB), GLM-4.5-Air in FP4 (~64 GB)gpt-oss-120b: ~21 GB for cache at ~0.04 MB per token. GLM-4.5-Air: ~120k tokens. A few to a dozen people
MoE ~235Bnot on one cardFP4 is ~141 GB2 cards in FP4, 4 cards in FP8
frontier 0.7-1T+not on one cardNVFP4 is ~400-415 GB8 cards (768 GB) in an RTX PRO Server, over PCIe

What you can actually do:

  • RAG for a team and a company: yes. A 30B model in FP8 or a 30B-A3B MoE serves dozens of people with a 16k-token context.
  • Agent backend for a company: yes, on 30B-120B models. gpt-oss-120b fits entirely on one card with a large context budget.
  • LoRA: on a BF16 base up to ~32B. QLoRA of a 70B model: the QLoRA authors fine-tuned a 65B model on a single 48 GB GPU.
  • Full fine-tuning: models of around 3-4B (96 GB ÷ 16 bytes per parameter is 6B before you count activations).
  • Many users: dozens on models up to 30B.
  • Several cards: 2 × 96 GB fits Qwen3-235B-A22B in FP4, 4 × Max-Q (384 GB) fits it in FP8 with ~110 GB for cache, and 8 cards in an RTX PRO Server (8 × 96 = 768 GB) fit ~700B models in NVFP4, such as DeepSeek-V3.1 (413 GB) or Mistral Large 3 (403 GB), with ~280 GB for cache. All of this, however, runs over PCIe, without NVLink.

B300: what does a single data-center GPU change?

288 GB of HBM3E and 8 TB/s, three times the memory and 4.5 times the bandwidth of the RTX PRO 6000. Plus NVLink 5 (1.8 TB/s per GPU) to the other GPUs in the node. Usable memory is about 259 GB.

Model classFits?PrecisionRealistic context and concurrency
dense 7-14ByesBF16over 1.4 million tokens of cache. Hundreds of conversations
dense 24-32ByesBF16 or FP8FP8: ~860k tokens in BF16
dense ~70ByesBF16 (~141 GB) or FP8 (~71 GB)FP8: ~575k tokens in BF16, ~1.15 million in FP8. For example 70 conversations of 16k each. Generation ceiling ~113 tokens/s per conversation
MoE ~30B-A3ByesBF16over 2 million tokens
MoE ~100-120ByesFP8 or FP4GLM-4.5-Air in FP8 (~113 GB): ~810k tokens
MoE ~235B (Qwen3-235B-A22B, DeepSeek-V4-Flash)yesFP4 (~141 GB) for Qwen3-235B. FP8 (236 GB) leaves ~24 GB, too little for many people. DeepSeek-V4-Flash: official files 167 GBQwen3-235B in FP4: ~610k tokens in BF16
frontier 0.7-1T+noNVFP4 is ~400 GB and upneeds at least 2 GPUs, in practice a node

What you can actually do:

  • Serving many users: this is what the tier is for. High bandwidth gives fast responses, and the large memory holds the cache of hundreds of conversations.
  • Agent backend for a company: models up to a 235B MoE on one GPU.
  • LoRA: on a BF16 base up to ~70B (141 GB of weights plus activations).
  • Full fine-tuning: ~8B models (16 bytes × 8.2B is about 131 GB plus activations), ~14B at the limit.
  • Training from scratch: small experimental models, not production-class ones (details in the cluster section).

The 8×B300 cluster: what fits on an HGX or DGX B300 node?

An HGX B300 or DGX B300 node is 8 Blackwell Ultra GPUs, 2.1 TB of HBM3E, NVLink 5 with switches (1.8 TB/s per GPU, 14.4 TB/s total), 108 PFLOPS FP4 (dense) and about 14 kW in a 10U chassis (DGX). Usable memory is about 1.9 TB. This is the first tier on which frontier-class open models fit entirely:

ModelParameters (total, active)Official filesOn 8×B300
DeepSeek-V3.1671B, 37B689 GB (FP8)yes, ~1.2 TB for cache
Mistral Large 3675B, 41B682 GB (FP8), 403 GB (NVFP4)yes. Mistral lists FP8 on a single B200 or H200 node
GLM-5744B, 40B756 GB (FP8 version)yes, ~1.1 TB for cache
Kimi K2.61T, 32B595 GB (native INT4)yes, ~1.3 TB for cache
DeepSeek-V4-Pro-08131.6T, 49B893 GByes. DeepSeek shows it running on a single 4×GB300 node
Kimi K32.8T, 104B1.56 TB (MXFP4)yes, ~330 GB for cache
Qwen3.8-2.4T-A95B2.4T, 95B2.5 TB (FP8 version)not in FP8. It needs FP4 quantization (~1.4-1.5 TB by our estimate) or two nodes

The DeepSeek-V4-Pro parameter counts come from the table in the DeepSeek-V4.1-Flash model card (1.6T "backbone" parameters, 49B active), because the V4-Pro-0813 card does not state them. File sizes are the sum of the safetensors files in the Hugging Face repositories.

The cache in these models is small thanks to MLA: for Kimi K2 and DeepSeek-V3 it is about 0.07 MB per token (from config.json: 61 layers × (512 + 64) × 2 bytes). With ~1.2 TB of free memory that means millions of cached tokens, or hundreds of parallel long-context agent conversations.

What you can actually do on 8×B300:

  • Serving frontier models to many users: yes, this is the main use.
  • LoRA and QLoRA of large MoE models: yes. OpenAI says gpt-oss-120b can be fine-tuned on a single H100 node (8 × 80 GB), and a B300 node has more than three times that memory.
  • Full fine-tuning: dense models up to about 70B (16 bytes × 70B is ~1.1 TB of model state, sharded across 8 GPUs with FSDP or ZeRO-3, plus activations).
  • Training from scratch: small models yes, frontier-level models no. The numbers are below.

How long would training from scratch take on one node?

Training cost is estimated with C ≈ 6 × N × D operations (N is parameters, D is training tokens) (Kaplan et al., 2020). HGX B300 lists 36 PFLOPS in BF16 with sparsity, which is 18 PFLOPS dense. We assume training uses 40% of that (our assumption, typical of well-optimized training, not a measurement), or ~7.2 PFLOPS.

Model and dataComputeTime on one 8×B300 (BF16, 40%)
1B on 20 billion tokens (~20 tokens per parameter, as in Chinchilla)1.2 × 10²⁰~4.6 hours
7B on 140 billion tokens5.9 × 10²¹~9.5 days
8B on 15 trillion tokens (current practice)7.2 × 10²³~3.2 years
70B on 15 trillion tokens6.3 × 10²⁴~28 years

Training in FP8 roughly doubles the pace, but the order of magnitude stays. For comparison with model cards: Llama 3.3 70B (trained on over 15 trillion tokens) used 7.0 million H100 GPU hours, DeepSeek-V3 2.788 million H800 GPU hours, and Mistral Large 3 was trained on 3,000 H200 GPUs. The conclusion: one node can train a small model from scratch for research or a narrow task, but a model that competes with the frontier is a job for clusters of thousands of GPUs on a fast network.

How splitting a model across several GPUs works (tensor parallelism and pipeline parallelism) is covered in Part 1. In short: PCIe 5.0 x16 gives 128 GB/s (NVIDIA, H100), while NVLink 5 on the B300 gives 1.8 TB/s per GPU, about 14 times more. The vLLM documentation recommends tensor parallelism when the model fits within one node, and pipeline parallelism for GPUs without NVLink (vLLM).

  • RTX PRO 6000: NVIDIA does not list NVLink in its specification. Scaling goes over PCIe. The 8-card RTX PRO Server uses ConnectX-8 cards with a built-in PCIe 6.0 switch, which NVIDIA says doubles inter-GPU bandwidth compared with discrete PCIe switches, but it is still not NVLink.
  • B300 in HGX/DGX: NVLink switches connect all 8 GPUs, so tensor parallelism across 8 GPUs is standard, and nodes connect over an 800 Gb/s network (ConnectX-8).

How much memory does fine-tuning need at these tiers?

The math is the same as in Part 1: QLoRA keeps the frozen base in 4 bits (Dettmers et al., 2023), LoRA in BF16 (Hu et al., 2021), and full fine-tuning with mixed-precision Adam needs 16 bytes per parameter, before activations (Rajbhandari et al., ZeRO).

TierQLoRALoRA (BF16 base)Full fine-tuning
RTX PRO 6000 96 GB~70Bup to ~32B~3-4B
B300 288 GBlarge MoE up to ~235Bup to ~70B~8B, 14B at the limit
8×B300 2.1 TBfrontier-class models~235B MoE in BF16dense up to ~70B

These are estimates from the formulas, rounded down, because activations grow with sequence length and batch size. Gradient checkpointing, 8-bit optimizers and offloading optimizer state to RAM push these limits up at the cost of time.

How do you choose between an RTX PRO 6000, a B300 and a cluster?

Start with the task, not the card: what model quality you need, how many people will use it at once, and how long a context you process.

Your situationSensible tier
RAG and agents for a company, dozens of users, data must stay in-houseRTX PRO 6000 Blackwell (1-4 cards)
A 235B or ~700B MoE on your own hardware without NVLink2-8 × RTX PRO 6000 with pipeline parallelism, with lower throughput than on B300
A production service for hundreds of users, 70-235B modelsB300 GPUs (owned or rented)
Frontier-class open models, full fine-tuning up to 70Ban 8×B300 node
Your own model, trained from scratch to compete with the frontiera multi-node cluster, in practice thousands of GPUs

If a personal assistant or RAG for a few people is enough, the 16 to 32 GB cards are covered in Part 1.

We show how we design AI on your own infrastructure, from choosing GPUs to serving models, on our AI infrastructure page. When it makes sense to move from an API to your own model at all, we cover in should you build or buy AI.

Sources

GPU specifications:

Model cards and configurations:

Method, serving and training: