What You Already Paid For and Haven't Switched On
A customer asked what database monitoring would cost. Ten minutes in it was clear they already had it. Entitled, not deployed, nobody had looked.
A customer asked what database monitoring would cost. Ten minutes in it was clear they already had it. Entitled, not deployed, nobody had looked.
Fifteen questions, about four minutes. Scores your estate across five areas and tells you what I'd do next in each one. Nothing is sent anywhere.
Five stages in the path, and only the last two involve the model. Retrieval is memory bound, generation is GPU bound, and they do not belong on the same hosts.
They implement the same ideas. The difference is what each optimises for and what it costs you to run, which is a less exciting answer than people want.
The arithmetic is easy. The denominator is where models fall apart, and the biggest lever has nothing to do with hardware.
Quantisation is usually framed as a quality tradeoff. The more useful framing is bandwidth, because decode re-reads every weight for every token. Tags: architecture, operations
Passthrough, time-sliced vGPU and MIG do different things to the thing that actually limits you. And fractional slices don't divide capacity the way the spreadsheet says.
Inference behaves like two workloads with opposite requirements, sharing the same silicon. Once you see that, the strange utilisation numbers start making sense.
Of all the decisions in a VCF build, tenancy is the one made worst. Not because it's hard, but because of when it gets made.
You upgraded and nothing changed. Most estates land on the platform and keep operating exactly as before. Here's the sequence I'd work through instead.
VMware on Azure, Google Cloud and Oracle has moved to BYOL, with deadlines. The nearest one is 31 October. And it quietly changes how you should compare clouds.
VCF on EC2 bare metal inside your own VPC. Not the managed service you may be thinking of, and the networking constraints will shape your design whether you like them or not.