Distributed Inference with llm-d and Kubernetes
Key Takeaways
Distributed inference with llm-d and Kubernetes for large language models
Original Description
As large language models move from research to running in production on Kubernetes, teams face the challenge of scaling inference efficiently, portably, and cost-effectively. llm-d, a new multi-vendor open source project re-imagines large-scale inference as a cloud-native system, integrating vLLM, the Kubernetes Gateway API Inference Extension, and new strategies such as prefill-decode disaggregation and precise KV-cache awareness.
This session introduces llm-d’s architecture and shows how it addresses key bottlenecks: GPU under-utilization, distributed KV-cache reuse, and complex request routing. Attendees will learn how they can achieve performant inference out of the box with llm-d, including features like prefill-decode disaggregation and distributed KV-cache reuse. Antonio Cardace (@acardace), Principal Machine Learning Engineer at Red Hat, is a dedicated technologist with a deep background in Kubernetes, Virtualization, and embedded development. An active member of the open-source community, he currently contributes to the llm-d project and previously served as a maintainer for the KubeVirt project.
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: LLM Engineering
View skill →Related Reads
📰
📰
📰
📰
this little its bitsy tiny gemma4 model on my 3060 is talking better than chat gpt
Reddit r/artificial
Prompt Injection for API Teams: What It Is and How to Test for It
Dev.to · Hassann
How to track new AI drops without the social media delay?
Reddit r/artificial
Inside the Model Factory — Eiso Kant, Poolside AI
Latent Space
🎓
Tutor Explanation
DeepCamp AI