A Quarter of the Card Is Not a Quarter of the Capacity
Passthrough, time-sliced vGPU and MIG do different things to the thing that actually limits you. And fractional slices don't divide capacity the way the spreadsheet says.
I ended the last piece on a point I want to pull apart properly, because it's the one I see costing people the most money.
Carving a GPU into fractional slices does not divide its usefulness proportionally. A quarter-card profile gives you considerably less than a quarter of the serving capacity, and the reason is arithmetic rather than overhead.
So here's how the three partitioning approaches actually work, what each does to the thing that matters, and how I'd pick between them.
Three ways to carve a GPU
Passthrough: the whole card, one VM
vGPU with time-slicing: sharing the clock
MIG: partitioning in hardware
The part that gets misjudged
Here's the arithmetic that catches people, and it follows directly from how inference uses memory.
Your serving capacity is set by the KV cache pool, which is whatever GPU memory remains once the model weights have loaded. Weights are a fixed cost. They don't shrink when you shrink the slice.
weights 16 GB → KV pool ~60 GB after overhead
Quarter profile, 20 GB
weights 16 GB → KV pool ~2 GB
# a quarter of the memory. Roughly 3% of the concurrency.
# illustrative numbers, but the shape is real.
Which gives you a rule worth remembering: fractional GPU profiles suit small models, not small workloads. If the model is large relative to the slice, fractioning doesn't help you serve more users, it stops you serving any.
Compute sharing versus memory partitioning
These get conflated constantly and they behave nothing alike.
That trade is the actual decision. Time-slicing gets you better aggregate utilisation and worse predictability. MIG gets you predictability and leaves capacity on the table when tenants are quiet.
The published figures from Broadcom's performance team showed utilisation climbing from around 37% with one tenant to 71% with four. That's the time-sliced behaviour working in your favour, with one tenant's idle compute absorbed by another. The data is here. Under strict hardware partitioning you wouldn't see the same curve, because the partitions can't lend each other anything.
How I'd actually choose
Use MIG when somebody has an SLO
Use time-sliced vGPU for development and mixed internal workloads
Use passthrough when the model needs the card
And check vMotion behaviour before you commit
The constraint most people meet late
Worth repeating from the multi-tenancy piece because it decides your design rather than just informing it: co-locating several tenants inside one VM is supported only with manual tenancy management. VCF Automation requires each tenant to have its own VM.
So if self-service is the goal, you're on VM-per-tenant, and your partitioning choice is between MIG-backed and time-sliced vGPU profiles. The denser container-based option is off the table until that changes.
Sizing a profile honestly
Where this is going
Two things worth watching. Quantisation changes the arithmetic above substantially, because smaller weights leave more room for cache. That's the next piece. And disaggregating prefill from decode would let you partition differently for each phase, which makes a lot more sense than treating them identically given how differently they behave.
None of this is specific to one hypervisor. MIG and vGPU are NVIDIA mechanisms, and the sizing arithmetic applies identically on bare metal, on Kubernetes with the device plugin, or anywhere else. What the platform decides is how cleanly you can present those partitions to tenants and whether your automation can drive them.
Sources
- NVIDIA Multi-Instance GPU user guide: partition geometry, supported profiles, and what is isolated.
- NVIDIA vGPU documentation: profiles, scheduling policies, and the mobility position for your driver version.
- Optimizing AI Deployments with VCF, Part 1, Broadcom: the multi-tenant utilisation figures and the VM co-location constraint.
- VCF and NVIDIA Hypervisor Certification.
- VCF on Supermicro HGX reference architecture: multi-GPU node design and fabric.
Profile availability, driver behaviour and mobility support all change between versions. Verify against the documentation for what you are actually running.
Views expressed here are my own.