> ## Content Index
> Fetch the complete content index at: https://www.mohammadsiddiqui.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Which Inference Engine, and Does It Matter?
- URL: https://www.mohammadsiddiqui.com/which-inference-engine-and-does-it-matter/
- Published: 2026-10-02T19:12:39.000Z
- Updated: 2026-10-02T19:12:39.000Z
- Description: They implement the same ideas. The difference is what each optimises for and what it costs you to run, which is a less exciting answer than people want.
- Author: Mohammad Siddiqui
- Tags: architecture, Ai, operations

Having written about what happens inside an inference engine, the obvious next question is which engine, and that turns out to be a less interesting question than people expect.

They mostly implement the same ideas. Paged KV cache, continuous batching, prefix reuse, the same quantisation formats. The differences are real but they're about what each one optimises for and what it costs you to run, rather than one being twice as fast as another in some general sense.

So this is less a league table and more a description of what you're choosing between, because the decision usually turns on operational fit rather than throughput.

What I'm not going to do

Publish benchmark numbers. Comparative inference benchmarks are enormously sensitive to model, hardware, sequence lengths, batch configuration and version, and most of the ones circulating are measuring something other than what they claim. If you want numbers for your decision, run them on your own workload. That's not a cop-out, it's the only method that produces a usable answer.

## The four you'll actually be choosing between

vLLM: the general-purpose default 

The project that introduced PagedAttention, and the one with the broadest model coverage and the widest hardware support. It has become the common substrate underneath a lot of platforms, including the VCF model runtime and several Kubernetes-based AI platforms. If you have no strong reason to pick something else, this is the sensible starting point, and the tuning knowledge transfers to wherever it shows up next. [The original paper](https://arxiv.org/abs/2309.06180?ref=mohammadsiddiqui.com) is still the clearest explanation of how the memory management works.

TensorRT-LLM: peak performance on NVIDIA, at a cost 

NVIDIA's own stack, built around compiled engines tuned for specific hardware. When it is the fastest option it tends to be meaningfully so, because the kernels are built for the silicon rather than being general. What you pay for that is a build step. You compile an engine for a given model, precision, batch and sequence configuration, and changing those usually means rebuilding. That's fine for a stable production workload and genuinely annoying during development.

SGLang: structured output and heavy prefix reuse 

Grew out of work on efficient prefix caching and structured generation. Where it tends to pull ahead is workloads with heavy prefix sharing or constrained output formats, which describes a lot of agentic and tool-calling work. If your application is doing JSON-shaped generation or running many requests that share a long prefix, it's worth a look.

TGI: simple to run 

Hugging Face's serving stack. Production-oriented, straightforward to operate, with good defaults. Tends to win on operational simplicity rather than peak throughput, which for plenty of teams is the right trade.

## Two things that confuse the picture

NIM is packaging, not an engine 

NVIDIA NIM bundles an engine, usually TensorRT-LLM, with a model, a container and an OpenAI-compatible API. So comparing NIM against vLLM is comparing a distribution against a component. What NIM is really selling is that somebody has already done the compilation and tuning for a given model and card, which is worth real money if that's the work you were dreading. [Documentation here](https://docs.nvidia.com/nim/?ref=mohammadsiddiqui.com).

Triton is the server, not the engine either 

Triton Inference Server handles the serving layer, model management and multi-model hosting, with the actual generation happening in a backend. You can run TensorRT-LLM or vLLM underneath it. Worth knowing because people compare it against engines and end up confused.

## What actually differs

Model coverage

How quickly a new architecture is supported. vLLM moves fast here because of the size of the community. A compiled stack usually lags because somebody has to do the kernel work.

Hardware breadth

Whether it runs on anything other than recent NVIDIA. Relevant if your estate has mixed accelerators or you want to avoid designing yourself into a corner.

Build and deployment friction

Does the model load and serve, or does it need compiling first? This is the biggest practical difference in day to day work and the one most likely to annoy your developers.

Prefix and cache behaviour

All of them do something here, but they differ in how aggressive and how configurable it is. Matters enormously if your prompts share long prefixes.

Quantisation formats

Which formats are supported and which are accelerated natively. Check this against the formats you actually intend to use rather than the list length.

Operational surface

Metrics, health, autoscaling signals, how it behaves when it runs out of KV cache. Unglamorous and it's what you'll live with.

## How I'd choose

Fewer questions than you'd think.

Is this a stable production workload with a fixed model, or will people be swapping models constantly? Stable favours a compiled stack, churn favours a flexible one Does your target model have good support in the stack today, not on a roadmap Will you ever run on anything other than current NVIDIA hardware Do your prompts share long prefixes, and is your output structured Who operates it at 2am, and does the stack give them anything useful when it misbehaves Can your team rebuild an engine if a model changes, or is that a blocker Have you run your own workload on the shortlist rather than reading someone else's benchmark

In practice I'd start on vLLM for almost anything, get the workload stable and understood, and only move to a compiled stack if the measurements say the performance gap is worth the operational friction. Starting with the most tuned option and working backwards tends to waste time, because you're optimising before you know what you're optimising.

## What doesn't change

Whichever you pick, the mechanics from [the first piece in this series](https://www.mohammadsiddiqui.com/your-gpu-is-not-slow-it-is-waiting-for-memory/) still apply. Prefill is compute-bound, decode is memory-bound, the KV cache is your capacity ceiling, and [quantisation](https://www.mohammadsiddiqui.com/smaller-numbers-faster-tokens/) moves the arithmetic. These are properties of transformer inference rather than of any particular implementation.

Which is the useful part for anyone in infrastructure. The engine is a choice you can revisit. The hardware, the partitioning model and the fabric are much harder to change, and they're decided by the same mechanics regardless of what you run on top.

## Sources

- [vLLM documentation](https://docs.vllm.ai/?ref=mohammadsiddiqui.com) and [the PagedAttention paper](https://arxiv.org/abs/2309.06180?ref=mohammadsiddiqui.com).
- [TensorRT-LLM documentation](https://nvidia.github.io/TensorRT-LLM/?ref=mohammadsiddiqui.com): the engine build workflow and supported configurations.
- [SGLang documentation](https://docs.sglang.ai/?ref=mohammadsiddiqui.com): prefix caching and structured generation.
- [Text Generation Inference documentation](https://huggingface.co/docs/text-generation-inference/?ref=mohammadsiddiqui.com), Hugging Face.
- [NVIDIA NIM documentation](https://docs.nvidia.com/nim/?ref=mohammadsiddiqui.com): what the packaging includes.

*All of these move quickly and capabilities converge. Anything specific here is worth rechecking against current documentation before it informs a decision.*

*Views expressed here are my own.*