Ai
RAG Is a Distributed Systems Problem
Five stages in the path, and only the last two involve the model. Retrieval is memory bound, generation is GPU bound, and they do not belong on the same hosts.
AI
Ai
Five stages in the path, and only the last two involve the model. Retrieval is memory bound, generation is GPU bound, and they do not belong on the same hosts.
architecture
They implement the same ideas. The difference is what each optimises for and what it costs you to run, which is a less exciting answer than people want.
architecture
The arithmetic is easy. The denominator is where models fall apart, and the biggest lever has nothing to do with hardware.
architecture
Quantisation is usually framed as a quality tradeoff. The more useful framing is bandwidth, because decode re-reads every weight for every token. Tags: architecture, operations
architecture
Passthrough, time-sliced vGPU and MIG do different things to the thing that actually limits you. And fractional slices don't divide capacity the way the spreadsheet says.
architecture
Inference behaves like two workloads with opposite requirements, sharing the same silicon. Once you see that, the strange utilisation numbers start making sense.
architecture
A single tenant on four GPUs used 37% of the streaming multiprocessors. Four tenants used 71%, with four times the throughput. Same box, no extra hardware.
security
A four-stage security framework with a scored assessment underneath it. The interesting part isn't the programme, it's the five pillars it scores you against, and the fact that one of them is people.
architecture
Two B200 nodes, one liquid-cooled and one air-cooled, 3.2 Tb/s of GPU fabric per node, and a four-host EPYC management domain. This is what the AI Factory announcement looks like when somebody actually builds it.
events
The hard part of enterprise AI was never the model. It's letting an agent use enterprise data without giving it access to enterprise data. Tanzu's answer is governed data products published to a marketplace. GA is Fall 2026.
events
An autonomous agent doesn't fit the user model or the service account model. Broadcom's answer gives agents identities, missions and runtime policy, and routes every action through a gateway that checks intent before execution.
News
Three announcements covering the platform, the automation underneath it, and the models that run on it. Here's the technical read and the questions I'd want answered before any of it lands in a design.