RAG Is a Distributed Systems Problem
Five stages in the path, and only the last two involve the model. Retrieval is memory bound, generation is GPU bound, and they do not belong on the same hosts.
Retrieval-augmented generation gets discussed as a model technique. Fetch some relevant documents, put them in the prompt, get a better answer.
From where I sit it's a distributed systems problem with a language model bolted on the end, and almost every RAG deployment I've seen struggle has struggled for infrastructure reasons rather than model ones.
So here's what's actually in the path, where the time goes, and the decisions that are expensive to reverse.
What a single request actually does
Five stages, and only the last two involve the model you were thinking about.
Everything before stage four is pure addition to time-to-first-token. The user is waiting and the GPU has not started.
The thing nobody budgets for
That matters for capacity planning, because the two phases want different things from the hardware. If you sized a GPU estate around a decode-heavy chat workload and then added RAG, your assumptions moved underneath you. The mechanics are in the first piece in this series if you want the detail.
It also means the retrieval decisions and the inference decisions aren't separable. How many chunks you retrieve, and how big they are, is a latency and cost decision as much as a quality one.
Prefix caching helps less than people hope
Here's a specific disappointment worth knowing in advance.
Prefix caching gives you a large win when many requests share the same long prefix, which is why it works so well for a chatbot with a fixed system prompt. RAG breaks that, because the retrieved context is different for every query. The shared part of your prompt is now just the system instructions, and the expensive part is unique.
You can claw some of it back by ordering the prompt so the stable parts come first and the retrieved chunks last, which at least lets the cache do something. Worth doing, and worth not expecting miracles from.
The index is a capacity planning exercise
Memory, roughly
Index type is a three-way trade
Embedding dimensions are a long-term commitment
Freshness is a pipeline, not a feature
Two workloads, not one
This is the architectural point I'd most want someone to take away.
Retrieval is memory and CPU bound. Generation is GPU bound. They scale on completely different curves, and if you co-locate them on the same nodes you end up buying GPUs to get more RAM, or RAM to get more GPU, depending on which way your load goes.
Separate them. Retrieval on memory-heavy hosts, generation on accelerated hosts, a network hop between. The hop costs you a few milliseconds and buys you the ability to scale each independently, which over the life of the platform is a much better trade.
The embedding model sits awkwardly between the two. It is a GPU workload but a small one, it is latency-sensitive because it is in the critical path, and it batches well. Worth deciding deliberately whether it shares accelerators with generation or gets its own, and the partitioning arithmetic applies to that decision in the usual way.
The hard problem nobody designs for
Pre-filtering also has a performance cost, because filtered vector search is harder than unfiltered. Plan for it rather than discovering it.
What I'd establish before building
A note on where this goes
Agentic systems make all of the above harder, because a single user request becomes several model calls with retrieval between them, and the latency budget that was tight for one round trip now has to cover five. If RAG is already marginal on latency, agentic workflows on the same infrastructure will not be.
Worth knowing before somebody demonstrates an agent and asks how quickly it can go into production.
Related
- Prefill, decode and the KV cache: why long prompts change which phase dominates.
- GPU partitioning: relevant to where the embedding model lives.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, Lewis et al. The original paper.
- Efficient and robust approximate nearest neighbor search using HNSW graphs, Malkov and Yashunin.
- vLLM documentation: prefix caching behaviour and configuration.
Vector database capabilities and index options vary considerably between products and change quickly. Check specifics against whatever you are actually running.
Views expressed here are my own.