Smaller Numbers, Faster Tokens
Quantisation is usually framed as a quality tradeoff. The more useful framing is bandwidth, because decode re-reads every weight for every token. Tags: architecture, operations
Quantisation gets talked about as a quality tradeoff. Shrink the numbers, lose a bit of accuracy, save some memory. That framing is true enough and it misses the thing that actually matters on an infrastructure sheet.
Smaller weights mean fewer bytes to move. And since decode is bound by memory bandwidth rather than compute, moving fewer bytes makes token generation faster. The memory saving is almost a side effect.
Once you see it that way the decision changes shape, because you're no longer trading quality for storage. You're trading quality for latency and concurrency, both of which somebody is probably complaining about.
This follows on from why decode is memory-bound in the first place and what fractional GPU profiles do to your capacity. You don't need either to read this, but the arithmetic at the end will make more sense if you've seen them.
What is actually being shrunk
Three different things can be quantised and they behave differently. Conflating them is where most confusion comes from.
Weights: the big one
Activations: harder, and often skipped
The KV cache: the one people forget
The formats, briefly
Where the speedup actually lands
This is the part that surprises people, and it comes straight from the two-phase behaviour.
So the honest answer to "how much faster will quantisation make this" is: depends entirely on your prompt-to-output ratio, which almost nobody measures before asking.
If you want both phases faster, you need the activations quantised too so the hardware can compute in the smaller format natively. That's a bigger step with more quality risk.
What it costs you in quality
Generalising here is dangerous, so I'll be careful.
Eight-bit, particularly FP8, is close enough to lossless for most production use that I'd treat it as the baseline rather than a compromise. Four-bit weight-only is usually a modest degradation on large models and a more noticeable one on small models, because smaller models have less redundancy to spare.
The important part is that degradation is not uniform across tasks. A model that still writes fluent prose after quantisation may be measurably worse at structured output, code, or arithmetic. So a general benchmark score is weak evidence for your specific use case.
What this does to your infrastructure arithmetic
This is where it connects back to partitioning, and it's the reason I wrote the two pieces in this order.
weights 16 GB → KV pool ~60 GB
Same card, 8-bit weights
weights 8 GB → KV pool ~68 GB
Same card, 4-bit weights
weights 4 GB → KV pool ~72 GB
# more room for context means more concurrent requests
# on hardware you already own. illustrative, but the direction is real.
And on a fractional profile the effect is much larger in proportion, because the weights were eating most of a small slice. Quantisation is what makes fractional GPU profiles viable for models that otherwise wouldn't fit usefully in them.
Which gives you a sequence worth following rather than treating these as independent choices. Quantise first, then size the profile, then decide the partitioning model. Doing it the other way round means sizing partitions for a footprint you were about to change.
What I'd do, in order
That last one matters more than it sounds. A deployment that was KV-cache-bound can become compute-bound after quantisation, and the right next action changes completely. If you're still tuning for the old constraint you'll be optimising something that stopped being the problem.
Vendor-neutral, as always
FP8, INT8, AWQ and GPTQ are not anybody's proprietary feature. They're supported across serving stacks, and the arithmetic above applies whether you're on VCF, on Kubernetes, or on a workstation under a desk. What differs is which formats your hardware accelerates natively, which is a silicon question rather than a platform one.
The platform decides how cleanly you can present the resulting partitions to tenants, and whether your automation can drive the lot. That's the layer I keep coming back to, because it's where the design work actually lives.
Sources
- vLLM quantisation documentation: supported formats and the configuration surface.
- AWQ: Activation-aware Weight Quantization, Lin et al.
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers, Frantar et al.
- FP8 Formats for Deep Learning: the format definitions and rationale.
- Optimizing AI Deployments with VCF, Part 1, Broadcom: GPU utilisation figures referenced above.
Quality impact is model and task specific, and format support changes with hardware and serving stack versions. Nothing here substitutes for testing against your own evaluation set.
Views expressed here are my own.