Which Inference Engine, and Does It Matter?

They implement the same ideas. The difference is what each optimises for and what it costs you to run, which is a less exciting answer than people want.

Share
Which Inference Engine, and Does It Matter?

Having written about what happens inside an inference engine, the obvious next question is which engine, and that turns out to be a less interesting question than people expect.

They mostly implement the same ideas. Paged KV cache, continuous batching, prefix reuse, the same quantisation formats. The differences are real but they're about what each one optimises for and what it costs you to run, rather than one being twice as fast as another in some general sense.

So this is less a league table and more a description of what you're choosing between, because the decision usually turns on operational fit rather than throughput.

What I'm not going to do
Publish benchmark numbers. Comparative inference benchmarks are enormously sensitive to model, hardware, sequence lengths, batch configuration and version, and most of the ones circulating are measuring something other than what they claim. If you want numbers for your decision, run them on your own workload. That's not a cop-out, it's the only method that produces a usable answer.

The four you'll actually be choosing between

vLLM: the general-purpose default
The project that introduced PagedAttention, and the one with the broadest model coverage and the widest hardware support. It has become the common substrate underneath a lot of platforms, including the VCF model runtime and several Kubernetes-based AI platforms. If you have no strong reason to pick something else, this is the sensible starting point, and the tuning knowledge transfers to wherever it shows up next. The original paper is still the clearest explanation of how the memory management works.
TensorRT-LLM: peak performance on NVIDIA, at a cost
NVIDIA's own stack, built around compiled engines tuned for specific hardware. When it is the fastest option it tends to be meaningfully so, because the kernels are built for the silicon rather than being general. What you pay for that is a build step. You compile an engine for a given model, precision, batch and sequence configuration, and changing those usually means rebuilding. That's fine for a stable production workload and genuinely annoying during development.
SGLang: structured output and heavy prefix reuse
Grew out of work on efficient prefix caching and structured generation. Where it tends to pull ahead is workloads with heavy prefix sharing or constrained output formats, which describes a lot of agentic and tool-calling work. If your application is doing JSON-shaped generation or running many requests that share a long prefix, it's worth a look.
TGI: simple to run
Hugging Face's serving stack. Production-oriented, straightforward to operate, with good defaults. Tends to win on operational simplicity rather than peak throughput, which for plenty of teams is the right trade.

Two things that confuse the picture

NIM is packaging, not an engine
NVIDIA NIM bundles an engine, usually TensorRT-LLM, with a model, a container and an OpenAI-compatible API. So comparing NIM against vLLM is comparing a distribution against a component. What NIM is really selling is that somebody has already done the compilation and tuning for a given model and card, which is worth real money if that's the work you were dreading. Documentation here.
Triton is the server, not the engine either
Triton Inference Server handles the serving layer, model management and multi-model hosting, with the actual generation happening in a backend. You can run TensorRT-LLM or vLLM underneath it. Worth knowing because people compare it against engines and end up confused.

What actually differs

Model coverage
How quickly a new architecture is supported. vLLM moves fast here because of the size of the community. A compiled stack usually lags because somebody has to do the kernel work.
Hardware breadth
Whether it runs on anything other than recent NVIDIA. Relevant if your estate has mixed accelerators or you want to avoid designing yourself into a corner.
Build and deployment friction
Does the model load and serve, or does it need compiling first? This is the biggest practical difference in day to day work and the one most likely to annoy your developers.
Prefix and cache behaviour
All of them do something here, but they differ in how aggressive and how configurable it is. Matters enormously if your prompts share long prefixes.
Quantisation formats
Which formats are supported and which are accelerated natively. Check this against the formats you actually intend to use rather than the list length.
Operational surface
Metrics, health, autoscaling signals, how it behaves when it runs out of KV cache. Unglamorous and it's what you'll live with.

How I'd choose

Fewer questions than you'd think.

In practice I'd start on vLLM for almost anything, get the workload stable and understood, and only move to a compiled stack if the measurements say the performance gap is worth the operational friction. Starting with the most tuned option and working backwards tends to waste time, because you're optimising before you know what you're optimising.

What doesn't change

Whichever you pick, the mechanics from the first piece in this series still apply. Prefill is compute-bound, decode is memory-bound, the KV cache is your capacity ceiling, and quantisation moves the arithmetic. These are properties of transformer inference rather than of any particular implementation.

Which is the useful part for anyone in infrastructure. The engine is a choice you can revisit. The hardware, the partitioning model and the fabric are much harder to change, and they're decided by the same mechanics regardless of what you run on top.

Sources

All of these move quickly and capabilities converge. Anything specific here is worth rechecking against current documentation before it informs a decision.

Views expressed here are my own.