Quick answer
The break-even point for your own server is monthly server cost divided by the blended API price per million tokens, calculated for your traffic. Two things decide it: how well you utilize the GPU and which API you compare against. Against expensive closed models, a single-card server pays off at a few hundred million tokens a month. Against cheap APIs, including the same open model bought from a hosting provider, break-even can sit above what one card can process at all. If data must not leave the company, cost stops being the only criterion.
All prices in this article are public list prices as of 8 October 2026, in USD, excluding VAT. We don't quote our own costs: the example rests entirely on explicit assumptions you can swap for your own.
What goes into the cost of a GPU server?
The card is only part of the bill. We calculate the monthly total cost of ownership like this:
monthly cost = (hardware ÷ months of use) + power + rack space + people's time
| Component | How to calculate | Watch out for |
|---|---|---|
| Hardware | GPU and server price ÷ useful life in months | Card prices rose sharply in 2026. NVIDIA's US Marketplace listed the RTX PRO 6000 Blackwell Workstation Edition at $16,000 in August, against $8,565 at launch in 2025 (Let's Data Science) |
| Power | average draw in kW × 730 hours × price per kWh | Count the whole server, not just the card, plus cooling |
| Rack space | colocation or a share of your own server room | Backup power, uplink, physical access |
| People | admin hours a month × hourly rate | Driver and engine updates, monitoring, hardware failures |
| Redundancy | a second card or a fallback API | Without it, one failed card stops the service |
People's time is the line most often left out. Serving engines ship new versions every few weeks, new models need updates, and GPUs under constant load do fail. Someone has to deal with all of that.
What do you actually pay per million API tokens?
API price lists show separate prices for input tokens (the prompt), output tokens (the answer) and input read from cache. To compare an API with a server, you need to blend them according to your traffic:
blended price = input share × (uncached part × input price + cached part × cache price) + output share × output price
The mix depends on the use case:
- RAG and knowledge bases: the prompt carries document excerpts, so input is many times longer than the answer.
- Agents: every step resends the whole history plus tool results, so input grows with each step. Much of it is a repeated prefix that hits the cache.
- Content generation: a short instruction and a long answer. Here the ratio flips and APIs are relatively expensive.
In our own agent and knowledge-base traffic, input tokens make up the vast majority. So the example below assumes 20 input tokens per output token, with half the input served from cache. That's an assumption, not a measurement of your system: the best approach is to measure your own traffic over a few days of a pilot.
Two traps:
- A token is not the same unit across providers. Anthropic says the tokenizer used since Claude 4.7 produces about 30% more tokens for the same text than the previous one (Anthropic pricing). When you compare with a local model, what matters is the cost of handling the same request, not the price per token.
- Cache writes cost money too. At Anthropic, writing to the 5-minute cache costs 1.25 times the input price. The example ignores this, which slightly understates the API cost.
What does a million API tokens cost in October 2026?
Standard prices per million tokens, and the blended price under our assumptions (20:1, half the input cached):
| Model | Input | Cached input | Output | Blended price |
|---|---|---|---|---|
| GPT-5.5 (OpenAI) | $5.00 | $0.50 | $30.00 | $4.05 |
| Claude Opus 5.5 (Anthropic) | $4.00 | $0.20 | $20.00 | $2.95 |
| Claude Sonnet 5.5 (Anthropic) | $2.00 | $0.10 | $10.00 | $1.48 |
| GPT-5.4-mini (OpenAI) | $0.75 | $0.075 | $4.50 | $0.61 |
| Gemini 3.8 Flash (Google, until 31 Dec 2026) | $0.75 | $0.075 | $3.75 | $0.57 |
| Qwen3.6-27B (DeepInfra, open model) | $0.32 | not listed | $3.20 | $0.46 |
| Claude Haiku 5.5 (Anthropic, prompts up to 100k) | $0.10 | $0.01 | $0.50 | $0.08 |
Sources: OpenAI, Anthropic, Google, DeepInfra. Google has announced that Gemini 3.8 Flash prices double from 1 January 2027, which pushes the blended rate to about $1.14. The Batch APIs at OpenAI and Anthropic cost half price if you can wait for results.
The row that matters most is Qwen3.6-27B. It's an open model you can run yourself, and the same model bought through an API costs a few tens of cents per million tokens. The fair cost comparison is exactly that: the same model on your hardware or at a provider. Comparing against closed models tells you more about the quality gap than about server economics.
Break-even: the formula and a worked example
break-even (tokens per month) = monthly server cost ÷ blended price per million tokens
The example's assumptions. Replace each one with your own:
| Assumption | Value in the example | Where it comes from |
|---|---|---|
| 96 GB-class GPU | $16,000 | RTX PRO 6000 Blackwell Workstation Edition on NVIDIA's US Marketplace, August 2026 (source) |
| Rest of the server (CPU, RAM, drives, chassis, PSUs) | $10,000 | assumption |
| Useful life | 36 months | assumption |
| Average server draw | 1 kW | assumption |
| Electricity price | $0.25 per kWh | assumption, use the rate on your bill |
| Colocation | $300 a month | assumption |
| Admin time | 8 h × $75 a month | assumption |
Result: hardware $722 + power $183 + colocation $300 + people $600 = about $1,805 a month.
Break-even at that cost:
| Compared with | Blended price | Break-even per month | Per day |
|---|---|---|---|
| GPT-5.5 | $4.05 | ~0.45 billion tokens | ~15 million |
| Claude Opus 5.5 | $2.95 | ~0.61 billion | ~20 million |
| Claude Sonnet 5.5 | $1.48 | ~1.22 billion | ~41 million |
| GPT-5.4-mini | $0.61 | ~2.97 billion | ~99 million |
| Gemini 3.8 Flash (2026 prices) | $0.57 | ~3.16 billion | ~105 million |
| Qwen3.6-27B via API | $0.46 | ~3.95 billion | ~132 million |
| Claude Haiku 5.5 | $0.08 | ~23.7 billion | ~790 million |
The spread is huge: from about fifteen million to several hundred million tokens a day. That's why break-even figures you find online differ by orders of magnitude. Each one compares against a different API and assumes different traffic.
Can one card actually process that many tokens?
Break-even only matters if the server can reach it. Throughput depends on the model, weight precision, prompt length and the number of concurrent requests, so it's best measured on your own traffic.
A public reference point: in a vLLM test on a single RTX PRO 6000 Blackwell Server Edition (50 concurrent requests, 100 input and 600 output tokens), a 32B model in BF16 reached about 970 tokens per second in total and Qwen3-14B about 2,030 (Database Mart). That's a test with short prompts and long answers. In RAG, where long prompts dominate and part of them hits the prefix cache, total tokens processed per second are usually higher, because the prompt is processed in parallel rather than token by token.
The arithmetic: 1,000 tokens per second for a whole month is about 2.5 billion tokens. But company traffic isn't flat around the clock: at night and at weekends the card sits idle. At 40% average utilization you're left with about a billion tokens a month.
What the example tells us:
- Against expensive closed models (GPT-5.5, Opus 5.5) break-even is within reach of one card at decent utilization.
- Against mid-tier models (Sonnet 5.5) you're close to the line, and the deciding factor is whether the local model gives you good enough quality.
- Against cheap APIs and the same open model from a provider one card won't reach break-even. Cost alone won't justify your own server.
The literature says the same: a 2025 cost analysis concludes that break-even depends on usage level and performance requirements, not on one magic number (Pan et al., arXiv 2509.18101).
Buy a server, rent a GPU or pay per token?
There's a third option: renting a cloud GPU by the hour. In October 2026 RunPod charges $2.09 an hour for an RTX PRO 6000 in Secure Cloud, $3.49 for an H100 SXM and $6.79 for a B200 (RunPod).
| Option | Monthly cost in the example | When it makes sense |
|---|---|---|
| Pay per token | depends on traffic | small or unpredictable traffic, project start, data may leave the company |
| Rent a GPU 24/7 | ~$1,526 + people's time | testing, pilots, seasonal traffic, no server room |
| Own server | ~$1,805 (including people's time) | steady traffic for years, data must stay in-house, your own models |
Buying for $26,000 pays back against 24/7 rental after about 25 months (26,000 ÷ (1,526 − 483), where $483 is power and colocation). With a shorter horizon, or if you only run during office hours, renting comes out cheaper. Keep in mind that a rented GPU is still someone else's infrastructure: check where it physically sits and who has access to it.
When is cost not the main argument?
Most companies we talk to about their own server don't start with cost. They start with whether data is allowed to leave the building. The arguments for your own server beyond price:
- Data stays on your network. Medical records, legal files, client data and trade secrets stay with you. We cover this in more detail in our article on closed models.
- Predictable cost. The server bill is fixed no matter how many times an agent calls the model.
- No rate limits. Your own server won't return a 429 in the middle of a batch job.
- Your own model. You serve a fine-tuned classifier or a LoRA adapter next to the base model, with no fees for hosting fine-tuned models. We walk through an example in small model with LoRA or big model with a prompt.
- A stable model version. The model on your server doesn't change unless you decide it should.
The arguments against deserve an honest hearing too:
- Quality. The best closed models are still sometimes better than open ones on hard tasks. How to pick an open model is the subject of Qwen, Gemma, GLM, Mistral or Bielik.
- Operations. A GPU server is one more production system with on-call duty, updates and failures.
- Flexibility. With an API, switching models is one line of code. On a server it means downloading weights, testing and deploying.
Checklist before you decide
- Measure your traffic. Input and output tokens per day, how much of each prompt repeats, and when the traffic happens.
- Calculate the blended price for two or three APIs, including the same open model at a hosting provider.
- Calculate the monthly server cost with every component, including people's time and redundancy.
- Measure throughput for your chosen model on your own traffic, not on a benchmark with short prompts.
- Assume realistic utilization, meaning your business hours, not 24/7.
- Check your data requirements: client contracts, professional secrecy, security policy. If data can't leave, the comparison with a public API falls away.
- Start with a pilot on a rented GPU or a demo before you buy hardware.
How we work it out ourselves
We run our own servers with NVIDIA Blackwell GPUs and serve models from the Qwen, Gemma, GLM and Mistral families on them. Before we tear down any model container, we record its counters for input and output tokens, prefix cache hits and preemptions. That way we know what real traffic looks like, not just a spreadsheet estimate. It's the same method we suggest to clients: measure on sample requests first, then decide on hardware.
If you want to see how many tokens your use case generates and how fast a model responds on a GPU, you can request a demo. How we design servers and model serving is described on our AI infrastructure page. We work out how much memory a model needs in our guide how much VRAM for a local LLM, and a broader comparison of deployment paths is in build or buy AI.
Sources
API price lists (as of 8 October 2026):
Hardware and GPU rental:
- Let's Data Science: Nvidia Lists RTX PRO 6000 Blackwell at $16,000 (12 Aug 2026)
- RunPod: Pricing
- Database Mart: vLLM GPU benchmark on the RTX PRO 6000
Method:
