What Does a Token Actually Cost You?
The arithmetic is easy. The denominator is where models fall apart, and the biggest lever has nothing to do with hardware.
Somebody will eventually ask what a token costs you. Usually it's finance, usually it's halfway through a project, and usually the answer is a shrug followed by a number nobody can defend.
The arithmetic isn't hard. What makes it difficult is that the denominator moves depending on decisions people haven't made yet, and most models quietly assume a throughput figure the system will never actually hit.
So here's how I'd build it, what breaks it, and the comparison against API pricing that the whole exercise usually exists to settle.
The model, at its simplest
That's it. Everything below is about getting those two numbers honest, and the second one is where people come unstuck.
The numerator: what an hour actually costs
Most models I've seen include the GPU and stop. The ones that survive a finance review include all of this.
The accelerator
The host around it
Power and cooling
Platform and licensing
Network and storage
People
The denominator: where models go wrong
Tokens per hour is not a property of your hardware. It's a property of your hardware under your workload, and the gap between those two is where most cost models fall apart.
That last one is worth sitting with, because it comes straight from the two-phase behaviour. Prefill processes your prompt in parallel and is compute-bound. Decode produces output one token at a time and is bound by memory bandwidth. They consume quite different resources per token.
Which is exactly why the API vendors price input and output tokens separately, usually with output several times more expensive. If your own model uses a single blended rate, it will be wrong in whichever direction your workload leans. A summarisation workload with long inputs and short outputs behaves nothing like a chat workload with the reverse.
Model them separately, or at minimum state your assumed ratio as an assumption rather than burying it.
Utilisation is the whole ballgame
Here is the thing that dominates everything else. The hardware costs the same whether it's busy or idle, so cost per token is really a function of how much of the hour you actually used.
Broadcom's performance team published figures in September showing a single tenant on a light inference workload using 36 to 38% of GPU streaming multiprocessors, and four tenants on the same hardware taking that to 71%, with roughly four times the throughput. The data is here.
Run that through the model and the cost per token in the four-tenant case is a fraction of the single-tenant case. Same hardware, same hourly cost, far more tokens out of it.
Comparing against an API
This is usually the real question behind the exercise, so let's be honest about it.
An API has no idle cost. You pay per token consumed, and if you use nothing you pay nothing. Owned infrastructure has a fixed cost that runs whether you use it or not. So there's a crossover volume below which the API wins on pure cost, and above which owning does.
Where that crossover sits depends entirely on your utilisation, which is why you can't answer the question without the number above.
When the API is genuinely the better answer
When owning starts to win
The reasons that have nothing to do with cost
What to hand finance
One page, with the assumptions visible rather than embedded in a spreadsheet nobody opens.
The sensitivities matter more than the headline figure. A model that only works at 70% utilisation is a model that depends on something you haven't built yet, and a good finance person will find that in about four minutes.
Quantisation moves the answer
Worth noting if you're building this now rather than later. Quantisation increases tokens per hour on the same hardware, so it reduces cost per token directly. If you're about to model this, model it on the configuration you intend to run rather than the one you're running today.
Same goes for consolidation. If the plan includes multi-tenancy, the cost model should reflect the consolidated utilisation rather than today's single-tenant number, with the consolidation plan attached as the thing that makes it real.
What I would actually do first
Measure utilisation. Not throughput, not latency, utilisation across a representative fortnight.
Most cost-per-token conversations I've been part of could have been resolved in an afternoon with that one number, because it usually turns out the hardware is fine and the way it's allocated is the problem.
Related
- Why decode is memory-bound, which is where the input versus output token asymmetry comes from.
- What fractional GPU profiles do to capacity, which affects the denominator more than people expect.
- Quantisation and what it buys you.
Sources
- Optimizing AI Deployments with VCF, Part 1, Broadcom: the utilisation and throughput figures used above.
- Private Cloud Outlook 2026, Broadcom: cost predictability as a repatriation driver. Vendor-commissioned survey, methodology disclosed.
Hardware pricing, instance rates and API pricing all move. Build the model so the inputs are easy to change, because you will be changing them.
Views expressed here are my own.