Smaller Numbers, Faster Tokens

Quantisation is usually framed as a quality tradeoff. The more useful framing is bandwidth, because decode re-reads every weight for every token. Tags: architecture, operations

Share
Smaller Numbers, Faster Tokens

Quantisation gets talked about as a quality tradeoff. Shrink the numbers, lose a bit of accuracy, save some memory. That framing is true enough and it misses the thing that actually matters on an infrastructure sheet.

Smaller weights mean fewer bytes to move. And since decode is bound by memory bandwidth rather than compute, moving fewer bytes makes token generation faster. The memory saving is almost a side effect.

Once you see it that way the decision changes shape, because you're no longer trading quality for storage. You're trading quality for latency and concurrency, both of which somebody is probably complaining about.

This follows on from why decode is memory-bound in the first place and what fractional GPU profiles do to your capacity. You don't need either to read this, but the arithmetic at the end will make more sense if you've seen them.

What is actually being shrunk

Three different things can be quantised and they behave differently. Conflating them is where most confusion comes from.

Weights: the big one
The model's parameters, loaded once and read on every forward pass. Dropping from 16-bit to 8-bit roughly halves both the memory they occupy and the bandwidth needed to read them. Four-bit halves it again. Since decode re-reads the weights for every single token, this goes straight to inter-token latency.
Activations: harder, and often skipped
The intermediate values flowing through the network. Quantising these as well as the weights is where you get compute savings, because the matrix units can then work in the smaller format natively. But activations have awkward outliers and are less forgiving than weights, which is why plenty of deployments quantise weights only and leave activations in 16-bit.
The KV cache: the one people forget
The cache can be stored in a smaller format too. This doesn't touch model quality in the same way, and it directly increases how many concurrent requests fit in a given amount of memory. If your deployment is KV-cache-bound rather than compute-bound, and most enterprise inference is, this is often the highest-value change available.

The formats, briefly

FP8
Eight-bit floating point, supported natively in hardware on recent data centre GPUs. Roughly half the memory and bandwidth of 16-bit, usually with very little quality impact. If your hardware supports it, this is the default starting point rather than an optimisation.
INT8
Eight-bit integer, typically applied to weights and activations together, and usually needing a calibration pass against representative data. More work than FP8 and more widely supported on older silicon.
Four-bit weight-only, AWQ and GPTQ
Two post-training approaches to compressing weights to four bits while keeping activations at higher precision. AWQ protects the weights that matter most to activation magnitudes. GPTQ works layer by layer, reconstructing to minimise error. Both give you a large memory and bandwidth win.

Where the speedup actually lands

This is the part that surprises people, and it comes straight from the two-phase behaviour.

Weight-only quantisation helps decode a lot and prefill very little
Decode is bandwidth-bound, so halving the bytes you read per token helps directly. Prefill is compute-bound, and if you have to unpack four-bit weights back to a higher precision to do the matrix maths, you've added work to a phase that was already compute-limited. On short prompts with long outputs you'll see a big win. On long prompts with short outputs you may see very little, and occasionally slightly worse.

So the honest answer to "how much faster will quantisation make this" is: depends entirely on your prompt-to-output ratio, which almost nobody measures before asking.

If you want both phases faster, you need the activations quantised too so the hardware can compute in the smaller format natively. That's a bigger step with more quality risk.

What it costs you in quality

Generalising here is dangerous, so I'll be careful.

Eight-bit, particularly FP8, is close enough to lossless for most production use that I'd treat it as the baseline rather than a compromise. Four-bit weight-only is usually a modest degradation on large models and a more noticeable one on small models, because smaller models have less redundancy to spare.

The important part is that degradation is not uniform across tasks. A model that still writes fluent prose after quantisation may be measurably worse at structured output, code, or arithmetic. So a general benchmark score is weak evidence for your specific use case.

The only test that counts
Run your own evaluation set, on your own prompts, before and after. If you don't have an evaluation set, that's the actual gap and it's worth more than any quantisation decision. Without one you have no way to know whether the thing you shipped got worse, which means you also can't tell when a model update breaks something.

What this does to your infrastructure arithmetic

This is where it connects back to partitioning, and it's the reason I wrote the two pieces in this order.

Whole 80 GB card, 16-bit weights
  weights 16 GB → KV pool ~60 GB

Same card, 8-bit weights
  weights 8 GB → KV pool ~68 GB

Same card, 4-bit weights
  weights 4 GB → KV pool ~72 GB

# more room for context means more concurrent requests
# on hardware you already own. illustrative, but the direction is real.

And on a fractional profile the effect is much larger in proportion, because the weights were eating most of a small slice. Quantisation is what makes fractional GPU profiles viable for models that otherwise wouldn't fit usefully in them.

Which gives you a sequence worth following rather than treating these as independent choices. Quantise first, then size the profile, then decide the partitioning model. Doing it the other way round means sizing partitions for a footprint you were about to change.

What I'd do, in order

That last one matters more than it sounds. A deployment that was KV-cache-bound can become compute-bound after quantisation, and the right next action changes completely. If you're still tuning for the old constraint you'll be optimising something that stopped being the problem.

Vendor-neutral, as always

FP8, INT8, AWQ and GPTQ are not anybody's proprietary feature. They're supported across serving stacks, and the arithmetic above applies whether you're on VCF, on Kubernetes, or on a workstation under a desk. What differs is which formats your hardware accelerates natively, which is a silicon question rather than a platform one.

The platform decides how cleanly you can present the resulting partitions to tenants, and whether your automation can drive the lot. That's the layer I keep coming back to, because it's where the design work actually lives.

THREE PARTS, IN ORDER
1
Your GPU is not slow, it is waiting for memory
prefill, decode and the KV cache
3
Smaller numbers, faster tokens  you are here
quantisation, and what it buys you

Sources

Quality impact is model and task specific, and format support changes with hardware and serving stack versions. Nothing here substitutes for testing against your own evaluation set.

Views expressed here are my own.