What Does a Token Actually Cost You?

The arithmetic is easy. The denominator is where models fall apart, and the biggest lever has nothing to do with hardware.

Share
What Does a Token Actually Cost You?

Somebody will eventually ask what a token costs you. Usually it's finance, usually it's halfway through a project, and usually the answer is a shrug followed by a number nobody can defend.

The arithmetic isn't hard. What makes it difficult is that the denominator moves depending on decisions people haven't made yet, and most models quietly assume a throughput figure the system will never actually hit.

So here's how I'd build it, what breaks it, and the comparison against API pricing that the whole exercise usually exists to settle.

The model, at its simplest

cost per token = cost per hour / tokens produced per hour

That's it. Everything below is about getting those two numbers honest, and the second one is where people come unstuck.

The numerator: what an hour actually costs

Most models I've seen include the GPU and stop. The ones that survive a finance review include all of this.

The accelerator
Either the amortised purchase cost over its useful life, or the instance hour if you're renting. Pick a life you can defend. Three years is conventional and probably optimistic given how fast this hardware is moving, five years is what a CFO will want, and the difference between them changes your answer substantially.
The host around it
CPU, memory, local storage, the chassis. A GPU server is not just GPUs, and on a decode-heavy workload the CPU is doing real work on tokenisation and scheduling. This gets dropped from models surprisingly often.
Power and cooling
Material at this density, and worth getting from whoever owns the facility rather than estimating. If you're liquid cooling, the capital side of that belongs here too.
Platform and licensing
Hypervisor, management, monitoring. Whatever your entitlement actually costs, apportioned across the estate. Easy to forget because it's paid centrally and doesn't appear on the project's own invoice.
Network and storage
Fabric for multi-node, and the storage holding model weights. Weight loading affects startup time rather than steady-state cost, but it matters if you're scaling instances up and down.
People
The one everybody leaves out, and often the largest line. Someone operates this. If you're comparing against an API where the vendor operates it, excluding your own operational effort makes the comparison meaningless.

The denominator: where models go wrong

Tokens per hour is not a property of your hardware. It's a property of your hardware under your workload, and the gap between those two is where most cost models fall apart.

The three mistakes I see most
Using peak benchmark throughput rather than throughput at your actual concurrency and sequence lengths. Ignoring idle time, because a GPU costs the same at 3am as it does at 11am. And counting input and output tokens as though they cost the same, which they don't.

That last one is worth sitting with, because it comes straight from the two-phase behaviour. Prefill processes your prompt in parallel and is compute-bound. Decode produces output one token at a time and is bound by memory bandwidth. They consume quite different resources per token.

Which is exactly why the API vendors price input and output tokens separately, usually with output several times more expensive. If your own model uses a single blended rate, it will be wrong in whichever direction your workload leans. A summarisation workload with long inputs and short outputs behaves nothing like a chat workload with the reverse.

Model them separately, or at minimum state your assumed ratio as an assumption rather than burying it.

Utilisation is the whole ballgame

Here is the thing that dominates everything else. The hardware costs the same whether it's busy or idle, so cost per token is really a function of how much of the hour you actually used.

Broadcom's performance team published figures in September showing a single tenant on a light inference workload using 36 to 38% of GPU streaming multiprocessors, and four tenants on the same hardware taking that to 71%, with roughly four times the throughput. The data is here.

Run that through the model and the cost per token in the four-tenant case is a fraction of the single-tenant case. Same hardware, same hourly cost, far more tokens out of it.

Which gives you the single most useful lever
Before optimising anything else, find out what your utilisation actually is. If a model shows your cost per token is uncompetitive, the problem is usually that you're paying for an idle accelerator rather than that the hardware is wrong. Consolidation fixes that and costs nothing.

Comparing against an API

This is usually the real question behind the exercise, so let's be honest about it.

An API has no idle cost. You pay per token consumed, and if you use nothing you pay nothing. Owned infrastructure has a fixed cost that runs whether you use it or not. So there's a crossover volume below which the API wins on pure cost, and above which owning does.

Where that crossover sits depends entirely on your utilisation, which is why you can't answer the question without the number above.

When the API is genuinely the better answer
Low or spiky volume. Early-stage projects where you don't yet know your usage pattern. Anything where you'd be buying hardware to sit idle most of the week. There's no shame in this and plenty of organisations should stay on APIs for longer than they do.
When owning starts to win
Sustained high volume, particularly with consolidation across several consumers. Predictable load. And the cases where cost isn't the deciding factor at all.
The reasons that have nothing to do with cost
Data that cannot leave your boundary. Residency obligations. Latency requirements that an internet round trip cannot meet. Cost predictability as a value in itself, which the Private Cloud Outlook survey found was climbing as a repatriation driver. If any of these apply, the cost comparison is informative rather than decisive, and you should say so out loud rather than torturing the model until it agrees with the decision you already made.

What to hand finance

One page, with the assumptions visible rather than embedded in a spreadsheet nobody opens.

The sensitivities matter more than the headline figure. A model that only works at 70% utilisation is a model that depends on something you haven't built yet, and a good finance person will find that in about four minutes.

Quantisation moves the answer

Worth noting if you're building this now rather than later. Quantisation increases tokens per hour on the same hardware, so it reduces cost per token directly. If you're about to model this, model it on the configuration you intend to run rather than the one you're running today.

Same goes for consolidation. If the plan includes multi-tenancy, the cost model should reflect the consolidated utilisation rather than today's single-tenant number, with the consolidation plan attached as the thing that makes it real.

What I would actually do first

Measure utilisation. Not throughput, not latency, utilisation across a representative fortnight.

Most cost-per-token conversations I've been part of could have been resolved in an afternoon with that one number, because it usually turns out the hardware is fine and the way it's allocated is the problem.

Related

Sources

Hardware pricing, instance rates and API pricing all move. Build the model so the inputs are easy to change, because you will be changing them.

Views expressed here are my own.