Your GPU Is Not Slow, It Is Waiting for Memory

Inference behaves like two workloads with opposite requirements, sharing the same silicon. Once you see that, the strange utilisation numbers start making sense.

Share
Your GPU Is Not Slow, It Is Waiting for Memory

Most infrastructure people I talk to about AI are reasoning about it the way we reason about any other workload. Allocate the GPU, size the memory, watch utilisation, add capacity when it fills up.

That model breaks almost immediately, and the reason is worth understanding properly rather than working around. Inference doesn't behave like one workload. It behaves like two, with opposite characteristics, sharing the same silicon and fighting each other for it.

Once you see that, a lot of things that look strange start making sense. Why a GPU reports 35% utilisation while refusing more requests. Why doubling batch size sometimes doubles throughput and sometimes destroys latency. Why the fractional GPU slice you carved up so neatly performs worse than the arithmetic suggested.

This is the layer above the hypervisor, and it's where I think infrastructure architects need to get to. None of it is VMware-specific, which is rather the point: the mechanics are the same whether you're running vLLM on VCF, on OpenShift, or on bare metal in someone else's data centre.

Two phases, pulling in opposite directions

When a request arrives, the model does two quite different jobs.

Prefill: reading the prompt
The model processes the entire prompt in one parallel pass and produces the first token. Attention across the prompt scales with the square of its length, so the arithmetic dominates. This phase is compute-bound. The GPU's matrix units are the constraint, and a long prompt is genuinely expensive. The metric people watch here is time to first token, TTFT.
Decode: writing the answer
Now the model emits tokens one at a time, each one depending on the last. The actual arithmetic per token is small. But to produce it, the GPU has to read the entire model's weights out of memory, plus the whole cached context so far. Every token. This phase is memory-bandwidth-bound, and the constraint is how fast you can move bytes out of HBM rather than how many FLOPS you have. The metric is time per output token, or inter-token latency.

Those two phases want opposite things from the scheduler. Prefill wants small batches so it doesn't blow out latency for everyone. Decode wants large batches, because if you're reading the weights out of memory anyway, you may as well amortise that read across as many sequences as possible.

Serving systems spend most of their effort managing that tension.

The one sentence to take away
Decode is memory-bound. That single fact explains most of what follows, including why GPU memory bandwidth and capacity matter more than headline FLOPS for a typical enterprise inference workload, and why your utilisation graphs look wrong.

The KV cache, and why it's the real capacity limit

To avoid recomputing attention over the whole sequence at every step, the model caches the key and value tensors it has already computed. That's the KV cache, and it's the central data structure in the whole business.

It grows with sequence length and with the number of concurrent requests. Roughly:

KV cache bytes = 2 × layers × kv_heads × head_dim × sequence_length × batch × bytes_per_element

# the 2 is keys and values. Grouped-query attention reduces kv_heads,
# which is why it changes your concurrency ceiling as much as it does.

Here's the consequence. After the model weights load, whatever GPU memory is left becomes the KV cache pool, and that pool is what decides how many requests you can hold in flight. Not compute. Memory.

Which means a GPU can be nowhere near compute-saturated and still refuse work, because it's out of somewhere to put the context. If you've seen a GPU sitting at modest utilisation while queueing requests, that's usually what you're looking at.

What PagedAttention fixed

Early serving systems allocated the KV cache as one contiguous block per sequence, sized for the longest output that request might produce. A request that finished after 200 tokens held memory reserved for thousands.

Published figures put the waste in those systems somewhere between 60 and 80 percent of KV cache memory.

vLLM borrowed virtual memory from operating systems. The cache is split into fixed-size blocks, a per-sequence block table maps logical token positions to physical blocks, and the blocks don't need to be contiguous. Waste drops to under 4 percent.

Why this matters to you specifically
It isn't just efficiency. Because blocks are independent, two requests with the same prefix can point at the same physical blocks. That's prefix caching, and in enterprise workloads it's a bigger win than it sounds, because so many requests share a long system prompt. Compute the prefix once, reuse it across every request that starts the same way. If your use case is a chatbot with a 2,000-token instruction block at the front of every call, prefix caching is doing a lot of quiet work for you.

Continuous batching, and the collision it creates

Static batching waits for a batch to fill, runs it, waits for the slowest request to finish, then starts again. The GPU idles while short requests sit finished alongside long ones.

Continuous batching schedules per iteration rather than per request. Finished sequences leave the batch, new ones join, every forward pass is full. Reported gains are in the region of two to three times the throughput of static batching.

But it creates a problem of its own. When a new request is admitted, its prefill is a big compute-bound operation that has to share a forward pass with everyone else's decode steps. One long prompt arriving can stall every token being generated for every other user.

Chunked prefill, the fix for that
Break the prefill into pieces and interleave them with decode steps rather than running the whole thing in one go. Reported improvements are in the region of 50 to 70 percent off p95 TTFT on mixed workloads. It's a flag rather than a redesign, and if your tail latency is bad on a workload with variable prompt lengths, it's the first thing I'd try.
Preemption, which you should know happens
When the KV pool fills, the scheduler has to evict something. Either it discards a request's blocks and recomputes them later, paying the prefill cost again, or it swaps them out. Either way somebody's request got slower for reasons that have nothing to do with their prompt. Worth knowing when you're staring at an unexplained latency spike.

Now the infrastructure part

This is where it connects to the decisions you and I actually make.

GPU selection: bandwidth over FLOPS, usually
If your workload is mostly decode, which most enterprise inference is, memory bandwidth and capacity determine throughput more than peak compute does. A card with more HBM and more bandwidth will serve more concurrent users than a card with better matrix throughput and less memory. Buying on FLOPS is buying for the prefill phase, which is not where your time goes.
Fractional GPU slices cut your KV pool, not just your compute
This is the one I see misjudged most, and I've since written up the partitioning arithmetic in full. Carving a GPU into fractional vGPU profiles partitions the framebuffer as well as the compute. A quarter-card profile gives you roughly a quarter of the KV cache pool, which means roughly a quarter of the concurrency, after the model weights have taken their cut off the top. And the weights don't shrink. So a model that leaves comfortable headroom on a whole card can leave almost none on a slice. The arithmetic is not linear and people are regularly surprised by it.
MIG versus time-slicing is a QoS decision
MIG partitions in hardware, giving each instance its own SM and memory slice. Predictable, isolated, and fixed once set. Time-sliced vGPU shares the GPU temporally, which is more elastic but means tenants can affect each other's latency. For a platform where several business units share accelerators and somebody has an SLO, MIG's predictability is usually worth the loss of flexibility. For a development estate, it usually isn't.
Why multi-tenancy improves utilisation, mechanically
Broadcom's performance team published figures in September showing a single tenant on a light inference workload using 36 to 38% of GPU streaming multiprocessors, rising to 71% with four tenants on the same hardware. Now you can see why. A light decode-heavy workload leaves the compute units idle while it waits on memory. Add more tenants and their work fills those gaps. You're not making any single workload faster. You're using the silicon that was already idle while one tenant waited for bytes to arrive. The data is here.
Interconnect matters once a model stops fitting
Tensor parallelism splits a model across GPUs and makes them talk constantly during every forward pass, so interconnect bandwidth becomes part of your inference latency. Within a node that's NVLink. Across nodes it's the fabric. The Supermicro HGX reference architecture published for VCF uses RoCEv2 Ethernet on Broadcom Tomahawk silicon rather than InfiniBand, with 3.2 Tb/s per node. Worth reading if you're designing anything multi-node, because the fabric decision is harder to reverse than the GPU one.
CPU topology still matters, which surprises people
GPU work doesn't exclude the CPU. Tokenisation, request scheduling, and the sampling loop all run there, and NUMA locality between the CPU handling a request and the GPU serving it affects latency. This is where the topology-aware scheduler work in 9.1 becomes relevant to AI workloads rather than just databases.
Memory tiering won't save you here
NVMe-backed memory tiering is a good answer to DRAM pricing for general workloads. It is not an answer for the KV cache, which needs to be read at HBM speeds on every decode step. Don't let the two conversations merge, because they will if nobody stops them.

Diagnosing a slow deployment

Three failure modes, and they want different fixes. This is the most practically useful thing in the whole piece.

GPU utilisation above 85%
You're genuinely compute-bound. More GPUs, a smaller model, or quantisation. This is the only one where buying hardware is the right answer.
High TTFT at moderate concurrency
KV-cache-bound. Raise the memory fraction given to the cache pool, enable chunked prefill, check whether quantisation buys you headroom. Adding GPUs may help, but tuning is cheaper and faster.
Utilisation under 60% at high concurrency
Scheduler-bound. Something upstream isn't feeding the GPU. Batching configuration, request routing, or the CPU side. Buying another accelerator here would be money set on fire.

If you take one operational habit from this: before anyone approves GPU spend, find out which of those three you're in. I'd guess most organisations have never checked.

Where this is heading

Two things worth watching, because they'll change how you design for this.

Disaggregated prefill and decode. Since the two phases want opposite hardware characteristics, run them on separate pools and ship the KV cache between them. Different instance types for different phases of the same request. It's an active area and it would reshape how you'd lay out a GPU estate.

KV cache as shared infrastructure. Caches that live outside a single serving instance, so a prefix computed on one node can be reused on another. That turns KV cache from a per-instance memory concern into something closer to a storage tier, with the naming, eviction and consistency problems that implies. Which is familiar territory for anyone who has designed a cache before.

None of this is VMware-specific

Everything above comes from how transformer inference works and how modern serving systems manage memory. It applies the same whether you're running vLLM under VCF, under OpenShift AI, under Ray, or on a bare metal box in a cupboard.

vLLM has become the common substrate across a lot of platforms, which is genuinely useful for us, because the tuning knowledge transfers. The alternatives are worth knowing too: TensorRT-LLM tends to win on NVIDIA hardware specifically, and SGLang and TGI have their own tradeoffs around memory efficiency and scheduling.

Where the platform does matter is everything underneath: how the GPU is presented, how tenants are isolated, what the fabric looks like, whether the CPU topology is working with you or against you. That's the layer I know best, and it's the one that decides whether the serving stack above it ever gets a chance to perform.

Things to go and check

THREE PARTS, IN ORDER
1
Your GPU is not slow, it is waiting for memory  you are here
prefill, decode and the KV cache
3
Smaller numbers, faster tokens
quantisation, and what it buys you

Sources

Serving systems move quickly and specific flags and defaults change between releases. The mechanics are stable, the tuning surface is not. Check the documentation for whatever version you're actually running.

Views expressed here are my own.