Quick answer

The break-even point for your own server is monthly server cost divided by the blended API price per million tokens, calculated for your traffic. Two things decide it: how well you utilize the GPU and which API you compare against. Against expensive closed models, a single-card server pays off at a few hundred million tokens a month. Against cheap APIs, including the same open model bought from a hosting provider, break-even can sit above what one card can process at all. If data must not leave the company, cost stops being the only criterion.

All prices in this article are public list prices as of 8 October 2026, in USD, excluding VAT. We don't quote our own costs: the example rests entirely on explicit assumptions you can swap for your own.

What goes into the cost of a GPU server?

The card is only part of the bill. We calculate the monthly total cost of ownership like this:

monthly cost = (hardware ÷ months of use) + power + rack space + people's time

ComponentHow to calculateWatch out for
HardwareGPU and server price ÷ useful life in monthsCard prices rose sharply in 2026. NVIDIA's US Marketplace listed the RTX PRO 6000 Blackwell Workstation Edition at $16,000 in August, against $8,565 at launch in 2025 (Let's Data Science)
Poweraverage draw in kW × 730 hours × price per kWhCount the whole server, not just the card, plus cooling
Rack spacecolocation or a share of your own server roomBackup power, uplink, physical access
Peopleadmin hours a month × hourly rateDriver and engine updates, monitoring, hardware failures
Redundancya second card or a fallback APIWithout it, one failed card stops the service

People's time is the line most often left out. Serving engines ship new versions every few weeks, new models need updates, and GPUs under constant load do fail. Someone has to deal with all of that.

What do you actually pay per million API tokens?

API price lists show separate prices for input tokens (the prompt), output tokens (the answer) and input read from cache. To compare an API with a server, you need to blend them according to your traffic:

blended price = input share × (uncached part × input price + cached part × cache price) + output share × output price

The mix depends on the use case:

  • RAG and knowledge bases: the prompt carries document excerpts, so input is many times longer than the answer.
  • Agents: every step resends the whole history plus tool results, so input grows with each step. Much of it is a repeated prefix that hits the cache.
  • Content generation: a short instruction and a long answer. Here the ratio flips and APIs are relatively expensive.

In our own agent and knowledge-base traffic, input tokens make up the vast majority. So the example below assumes 20 input tokens per output token, with half the input served from cache. That's an assumption, not a measurement of your system: the best approach is to measure your own traffic over a few days of a pilot.

Two traps:

  1. A token is not the same unit across providers. Anthropic says the tokenizer used since Claude 4.7 produces about 30% more tokens for the same text than the previous one (Anthropic pricing). When you compare with a local model, what matters is the cost of handling the same request, not the price per token.
  2. Cache writes cost money too. At Anthropic, writing to the 5-minute cache costs 1.25 times the input price. The example ignores this, which slightly understates the API cost.

What does a million API tokens cost in October 2026?

Standard prices per million tokens, and the blended price under our assumptions (20:1, half the input cached):

ModelInputCached inputOutputBlended price
GPT-5.5 (OpenAI)$5.00$0.50$30.00$4.05
Claude Opus 5.5 (Anthropic)$4.00$0.20$20.00$2.95
Claude Sonnet 5.5 (Anthropic)$2.00$0.10$10.00$1.48
GPT-5.4-mini (OpenAI)$0.75$0.075$4.50$0.61
Gemini 3.8 Flash (Google, until 31 Dec 2026)$0.75$0.075$3.75$0.57
Qwen3.6-27B (DeepInfra, open model)$0.32not listed$3.20$0.46
Claude Haiku 5.5 (Anthropic, prompts up to 100k)$0.10$0.01$0.50$0.08

Sources: OpenAI, Anthropic, Google, DeepInfra. Google has announced that Gemini 3.8 Flash prices double from 1 January 2027, which pushes the blended rate to about $1.14. The Batch APIs at OpenAI and Anthropic cost half price if you can wait for results.

The row that matters most is Qwen3.6-27B. It's an open model you can run yourself, and the same model bought through an API costs a few tens of cents per million tokens. The fair cost comparison is exactly that: the same model on your hardware or at a provider. Comparing against closed models tells you more about the quality gap than about server economics.

Break-even: the formula and a worked example

break-even (tokens per month) = monthly server cost ÷ blended price per million tokens

The example's assumptions. Replace each one with your own:

AssumptionValue in the exampleWhere it comes from
96 GB-class GPU$16,000RTX PRO 6000 Blackwell Workstation Edition on NVIDIA's US Marketplace, August 2026 (source)
Rest of the server (CPU, RAM, drives, chassis, PSUs)$10,000assumption
Useful life36 monthsassumption
Average server draw1 kWassumption
Electricity price$0.25 per kWhassumption, use the rate on your bill
Colocation$300 a monthassumption
Admin time8 h × $75 a monthassumption

Result: hardware $722 + power $183 + colocation $300 + people $600 = about $1,805 a month.

Break-even at that cost:

Compared withBlended priceBreak-even per monthPer day
GPT-5.5$4.05~0.45 billion tokens~15 million
Claude Opus 5.5$2.95~0.61 billion~20 million
Claude Sonnet 5.5$1.48~1.22 billion~41 million
GPT-5.4-mini$0.61~2.97 billion~99 million
Gemini 3.8 Flash (2026 prices)$0.57~3.16 billion~105 million
Qwen3.6-27B via API$0.46~3.95 billion~132 million
Claude Haiku 5.5$0.08~23.7 billion~790 million

The spread is huge: from about fifteen million to several hundred million tokens a day. That's why break-even figures you find online differ by orders of magnitude. Each one compares against a different API and assumes different traffic.

Can one card actually process that many tokens?

Break-even only matters if the server can reach it. Throughput depends on the model, weight precision, prompt length and the number of concurrent requests, so it's best measured on your own traffic.

A public reference point: in a vLLM test on a single RTX PRO 6000 Blackwell Server Edition (50 concurrent requests, 100 input and 600 output tokens), a 32B model in BF16 reached about 970 tokens per second in total and Qwen3-14B about 2,030 (Database Mart). That's a test with short prompts and long answers. In RAG, where long prompts dominate and part of them hits the prefix cache, total tokens processed per second are usually higher, because the prompt is processed in parallel rather than token by token.

The arithmetic: 1,000 tokens per second for a whole month is about 2.5 billion tokens. But company traffic isn't flat around the clock: at night and at weekends the card sits idle. At 40% average utilization you're left with about a billion tokens a month.

What the example tells us:

  • Against expensive closed models (GPT-5.5, Opus 5.5) break-even is within reach of one card at decent utilization.
  • Against mid-tier models (Sonnet 5.5) you're close to the line, and the deciding factor is whether the local model gives you good enough quality.
  • Against cheap APIs and the same open model from a provider one card won't reach break-even. Cost alone won't justify your own server.

The literature says the same: a 2025 cost analysis concludes that break-even depends on usage level and performance requirements, not on one magic number (Pan et al., arXiv 2509.18101).

Buy a server, rent a GPU or pay per token?

There's a third option: renting a cloud GPU by the hour. In October 2026 RunPod charges $2.09 an hour for an RTX PRO 6000 in Secure Cloud, $3.49 for an H100 SXM and $6.79 for a B200 (RunPod).

OptionMonthly cost in the exampleWhen it makes sense
Pay per tokendepends on trafficsmall or unpredictable traffic, project start, data may leave the company
Rent a GPU 24/7~$1,526 + people's timetesting, pilots, seasonal traffic, no server room
Own server~$1,805 (including people's time)steady traffic for years, data must stay in-house, your own models

Buying for $26,000 pays back against 24/7 rental after about 25 months (26,000 ÷ (1,526 − 483), where $483 is power and colocation). With a shorter horizon, or if you only run during office hours, renting comes out cheaper. Keep in mind that a rented GPU is still someone else's infrastructure: check where it physically sits and who has access to it.

When is cost not the main argument?

Most companies we talk to about their own server don't start with cost. They start with whether data is allowed to leave the building. The arguments for your own server beyond price:

  • Data stays on your network. Medical records, legal files, client data and trade secrets stay with you. We cover this in more detail in our article on closed models.
  • Predictable cost. The server bill is fixed no matter how many times an agent calls the model.
  • No rate limits. Your own server won't return a 429 in the middle of a batch job.
  • Your own model. You serve a fine-tuned classifier or a LoRA adapter next to the base model, with no fees for hosting fine-tuned models. We walk through an example in small model with LoRA or big model with a prompt.
  • A stable model version. The model on your server doesn't change unless you decide it should.

The arguments against deserve an honest hearing too:

  • Quality. The best closed models are still sometimes better than open ones on hard tasks. How to pick an open model is the subject of Qwen, Gemma, GLM, Mistral or Bielik.
  • Operations. A GPU server is one more production system with on-call duty, updates and failures.
  • Flexibility. With an API, switching models is one line of code. On a server it means downloading weights, testing and deploying.

Checklist before you decide

  1. Measure your traffic. Input and output tokens per day, how much of each prompt repeats, and when the traffic happens.
  2. Calculate the blended price for two or three APIs, including the same open model at a hosting provider.
  3. Calculate the monthly server cost with every component, including people's time and redundancy.
  4. Measure throughput for your chosen model on your own traffic, not on a benchmark with short prompts.
  5. Assume realistic utilization, meaning your business hours, not 24/7.
  6. Check your data requirements: client contracts, professional secrecy, security policy. If data can't leave, the comparison with a public API falls away.
  7. Start with a pilot on a rented GPU or a demo before you buy hardware.

How we work it out ourselves

We run our own servers with NVIDIA Blackwell GPUs and serve models from the Qwen, Gemma, GLM and Mistral families on them. Before we tear down any model container, we record its counters for input and output tokens, prefix cache hits and preemptions. That way we know what real traffic looks like, not just a spreadsheet estimate. It's the same method we suggest to clients: measure on sample requests first, then decide on hardware.

If you want to see how many tokens your use case generates and how fast a model responds on a GPU, you can request a demo. How we design servers and model serving is described on our AI infrastructure page. We work out how much memory a model needs in our guide how much VRAM for a local LLM, and a broader comparison of deployment paths is in build or buy AI.

Sources

API price lists (as of 8 October 2026):

Hardware and GPU rental:

Method: