About

Raghav Potluri

I'm a senior principal engineer working on the infrastructure under large language models — the parts that decide where a request goes, what gets cached, and what you can see when it's slow. Distributed inference, prompt caching, KV-cache coordination, tokenization cost, and telemetry. Right now I'm focused on AI inference routing. I'm an IEEE Senior Member and a Fellow of BCS, The Chartered Institute for IT, and I serve on the board of Educate2Envision, an education nonprofit.

I write here as InfraWhisperer. The posts are field notes: read the code, run it on real GPUs, report what actually happens. I spend most of my time in NVIDIA Dynamo, llm-d, vLLM, and SGLang — disaggregated prefill/decode, routing, and the KV-cache event plane that ties them together.

No vendor pitch. Views are my own. When I'm wrong, I say so.