A Quarter of the Card Is Not a Quarter of the Capacity

Passthrough, time-sliced vGPU and MIG do different things to the thing that actually limits you. And fractional slices don't divide capacity the way the spreadsheet says.

Share
A Quarter of the Card Is Not a Quarter of the Capacity

I ended the last piece on a point I want to pull apart properly, because it's the one I see costing people the most money.

Carving a GPU into fractional slices does not divide its usefulness proportionally. A quarter-card profile gives you considerably less than a quarter of the serving capacity, and the reason is arithmetic rather than overhead.

So here's how the three partitioning approaches actually work, what each does to the thing that matters, and how I'd pick between them.

Three ways to carve a GPU

Passthrough: the whole card, one VM
The GPU is handed directly to a virtual machine. Nothing sits in the path, so performance is as close to bare metal as you'll get, and the driver stack inside the guest is the normal one. The cost is flexibility. One VM, one card, and the mobility options are more limited than people expect. Fine for a dedicated training box or a single large model that needs everything. Wasteful for the chatbot nobody uses after 6pm.
vGPU with time-slicing: sharing the clock
The hypervisor presents virtual GPUs to several VMs and context-switches between them. Each vGPU gets a slice of the framebuffer, fixed by the profile you choose, and takes turns on the compute units. Scheduling policy decides how those turns are allocated: best effort, equal share, or fixed share. The important detail is that memory is partitioned and compute is shared. Those behave very differently.
MIG: partitioning in hardware
Multi-Instance GPU splits a supported card into instances that each get their own slice of SMs, their own L2 cache slice, and their own memory. Not a time slice. A physical partition, enforced below the driver. Each instance behaves like a smaller, independent GPU. You set the geometry up front and changing it means tearing instances down.

The part that gets misjudged

Here's the arithmetic that catches people, and it follows directly from how inference uses memory.

Your serving capacity is set by the KV cache pool, which is whatever GPU memory remains once the model weights have loaded. Weights are a fixed cost. They don't shrink when you shrink the slice.

Whole 80 GB card
  weights 16 GB → KV pool ~60 GB after overhead

Quarter profile, 20 GB
  weights 16 GB → KV pool ~2 GB

# a quarter of the memory. Roughly 3% of the concurrency.
# illustrative numbers, but the shape is real.
Why this bites
Halve the slice and you more than halve the concurrency. Quarter it and you may have nothing useful left at all, because the weights consume the same amount regardless. Anyone sizing fractional profiles on a spreadsheet that divides capacity linearly is going to be wrong, and the smaller the slice the more wrong they'll be.

Which gives you a rule worth remembering: fractional GPU profiles suit small models, not small workloads. If the model is large relative to the slice, fractioning doesn't help you serve more users, it stops you serving any.

Compute sharing versus memory partitioning

These get conflated constantly and they behave nothing alike.

Framebuffer is partitioned, always
Under vGPU each profile owns its memory allocation. It cannot borrow from a neighbour, even if that neighbour is idle. This is the hard ceiling on how much context you can hold.
Compute is shared under time-slicing
Everyone takes turns on the same SMs. A tenant running something heavy affects the latency everyone else sees. It also means an idle tenant's compute is available to others, which is why utilisation goes up when you add tenants.
MIG partitions both
SM slices and memory slices, fixed in hardware. Nobody steals anybody's compute, and nobody benefits from anybody's idleness either. Predictable, and less efficient at low load.

That trade is the actual decision. Time-slicing gets you better aggregate utilisation and worse predictability. MIG gets you predictability and leaves capacity on the table when tenants are quiet.

The published figures from Broadcom's performance team showed utilisation climbing from around 37% with one tenant to 71% with four. That's the time-sliced behaviour working in your favour, with one tenant's idle compute absorbed by another. The data is here. Under strict hardware partitioning you wouldn't see the same curve, because the partitions can't lend each other anything.

How I'd actually choose

Use MIG when somebody has an SLO
If different business units share accelerators and one of them has a latency commitment to defend, hardware isolation is worth what it costs. A noisy neighbour cannot reach across a MIG boundary. That argument also lands better with a risk function than 'we use a fair scheduler', because it's a structural guarantee rather than a policy.
Use time-sliced vGPU for development and mixed internal workloads
Where the tenants are internal, nobody has signed up to a latency number, and the goal is to get more out of what you own. The elasticity is genuinely useful and the occasional interference is tolerable.
Use passthrough when the model needs the card
Large model, tensor parallelism across several GPUs, or a training workload. If you're going to use the whole thing anyway, don't add a layer for nothing.
And check vMotion behaviour before you commit
Mobility support differs across passthrough, vGPU and MIG-backed vGPU, and it has changed between driver versions. It also affects whether a host can be put into maintenance mode without evicting workloads. I'd verify the current position against the NVIDIA vGPU documentation for the exact driver and vSphere version you're running rather than trusting any blog post, including this one.

The constraint most people meet late

Worth repeating from the multi-tenancy piece because it decides your design rather than just informing it: co-locating several tenants inside one VM is supported only with manual tenancy management. VCF Automation requires each tenant to have its own VM.

So if self-service is the goal, you're on VM-per-tenant, and your partitioning choice is between MIG-backed and time-sliced vGPU profiles. The denser container-based option is off the table until that changes.

Sizing a profile honestly

Where this is going

Two things worth watching. Quantisation changes the arithmetic above substantially, because smaller weights leave more room for cache. That's the next piece. And disaggregating prefill from decode would let you partition differently for each phase, which makes a lot more sense than treating them identically given how differently they behave.

None of this is specific to one hypervisor. MIG and vGPU are NVIDIA mechanisms, and the sizing arithmetic applies identically on bare metal, on Kubernetes with the device plugin, or anywhere else. What the platform decides is how cleanly you can present those partitions to tenants and whether your automation can drive them.

THREE PARTS, IN ORDER
1
Your GPU is not slow, it is waiting for memory
prefill, decode and the KV cache
2
A quarter of the card is not a quarter of the capacity  you are here
passthrough, vGPU and MIG
3
Smaller numbers, faster tokens
quantisation, and what it buys you

Sources

Profile availability, driver behaviour and mobility support all change between versions. Verify against the documentation for what you are actually running.

Views expressed here are my own.