RAG Is a Distributed Systems Problem

Five stages in the path, and only the last two involve the model. Retrieval is memory bound, generation is GPU bound, and they do not belong on the same hosts.

Share
RAG Is a Distributed Systems Problem

Retrieval-augmented generation gets discussed as a model technique. Fetch some relevant documents, put them in the prompt, get a better answer.

From where I sit it's a distributed systems problem with a language model bolted on the end, and almost every RAG deployment I've seen struggle has struggled for infrastructure reasons rather than model ones.

So here's what's actually in the path, where the time goes, and the decisions that are expensive to reverse.

What a single request actually does

Five stages, and only the last two involve the model you were thinking about.

1. Embed the query
Run the user's question through an embedding model to get a vector. A small GPU or CPU workload, but it is in the critical path and it is per request.
2. Search the index
Find the nearest vectors. Memory-bound and usually CPU-bound, with latency depending heavily on index type and size.
3. Rerank, if you do it
A second, usually larger model scores the candidates properly. Improves quality noticeably and adds latency just as noticeably.
4. Prefill
The retrieved context goes into the prompt, and now your prompt is long. This is the stage people forget about.
5. Decode
Generate the answer, one token at a time, as usual.

Everything before stage four is pure addition to time-to-first-token. The user is waiting and the GPU has not started.

The thing nobody budgets for

Retrieval makes your prompts long, and prefill is compute-bound
A bare chat request might have a few hundred tokens of prompt. The same request with six retrieved chunks might have four thousand. Attention cost in prefill scales with the square of prompt length, so you have not added a bit of latency, you have changed which phase dominates your request. A workload that was comfortably decode-bound can become prefill-heavy the day somebody turns on retrieval.

That matters for capacity planning, because the two phases want different things from the hardware. If you sized a GPU estate around a decode-heavy chat workload and then added RAG, your assumptions moved underneath you. The mechanics are in the first piece in this series if you want the detail.

It also means the retrieval decisions and the inference decisions aren't separable. How many chunks you retrieve, and how big they are, is a latency and cost decision as much as a quality one.

Prefix caching helps less than people hope

Here's a specific disappointment worth knowing in advance.

Prefix caching gives you a large win when many requests share the same long prefix, which is why it works so well for a chatbot with a fixed system prompt. RAG breaks that, because the retrieved context is different for every query. The shared part of your prompt is now just the system instructions, and the expensive part is unique.

You can claw some of it back by ordering the prompt so the stable parts come first and the retrieved chunks last, which at least lets the cache do something. Worth doing, and worth not expecting miracles from.

The index is a capacity planning exercise

Memory, roughly
A vector index is mostly a large array of floats plus a graph structure for searching it. The vectors themselves are approximately the number of documents times the embedding dimensions times four bytes at full precision, and a graph-based index adds meaningful overhead on top. Work that out before choosing an embedding model, because dimension count has a direct and permanent effect on your memory bill.
Index type is a three-way trade
Graph-based indexes like HNSW give excellent recall and low latency and eat memory. Quantised or partitioned indexes use far less memory and give up some recall. Flat search is exact and does not scale. Pick deliberately, because migrating between index types on a large corpus is a rebuild rather than a setting. The HNSW paper is worth reading if you want to understand why the memory goes where it does.
Embedding dimensions are a long-term commitment
Changing embedding model means re-embedding the entire corpus and rebuilding the index. On a large document set that is days of compute and a cutover plan. Choose as though you will live with it, because you will.
Freshness is a pipeline, not a feature
Documents change. Something has to notice, re-chunk, re-embed and update the index, and that something runs continuously rather than at setup. The staleness tolerance you can live with decides whether this is a nightly batch or a streaming pipeline, and those are very different builds.

Two workloads, not one

This is the architectural point I'd most want someone to take away.

Retrieval is memory and CPU bound. Generation is GPU bound. They scale on completely different curves, and if you co-locate them on the same nodes you end up buying GPUs to get more RAM, or RAM to get more GPU, depending on which way your load goes.

Separate them. Retrieval on memory-heavy hosts, generation on accelerated hosts, a network hop between. The hop costs you a few milliseconds and buys you the ability to scale each independently, which over the life of the platform is a much better trade.

The embedding model sits awkwardly between the two. It is a GPU workload but a small one, it is latency-sensitive because it is in the critical path, and it batches well. Worth deciding deliberately whether it shares accelerators with generation or gets its own, and the partitioning arithmetic applies to that decision in the usual way.

The hard problem nobody designs for

Retrieval has to respect document permissions
If your corpus contains documents with different access rights, and in an enterprise it does, then retrieval must filter by what the asking user is allowed to see. Not the generation step. Retrieval. By the time a document is in the prompt you have already leaked it, and no amount of prompting the model to be careful fixes that. This has to be designed in at the index level, with permissions either stored as metadata and applied as a pre-filter or enforced by partitioning the index. Retrofitting it later is unpleasant, and it is the failure mode most likely to end up in an incident report.

Pre-filtering also has a performance cost, because filtered vector search is harder than unfiltered. Plan for it rather than discovering it.

What I'd establish before building

A note on where this goes

Agentic systems make all of the above harder, because a single user request becomes several model calls with retrieval between them, and the latency budget that was tight for one round trip now has to cover five. If RAG is already marginal on latency, agentic workflows on the same infrastructure will not be.

Worth knowing before somebody demonstrates an agent and asks how quickly it can go into production.

Related

Sources

Vector database capabilities and index options vary considerably between products and change quickly. Check specifics against whatever you are actually running.

Views expressed here are my own.