Distributed Inference with llm-d and Kubernetes

Siemens Knowledge Hub · Beginner ·🧠 Large Language Models ·1mo ago

Key Takeaways

Distributed inference with llm-d and Kubernetes for large language models

Original Description

As large language models move from research to running in production on Kubernetes, teams face the challenge of scaling inference efficiently, portably, and cost-effectively. llm-d, a new multi-vendor open source project re-imagines large-scale inference as a cloud-native system, integrating vLLM, the Kubernetes Gateway API Inference Extension, and new strategies such as prefill-decode disaggregation and precise KV-cache awareness. This session introduces llm-d’s architecture and shows how it addresses key bottlenecks: GPU under-utilization, distributed KV-cache reuse, and complex request routing. Attendees will learn how they can achieve performant inference out of the box with llm-d, including features like prefill-decode disaggregation and distributed KV-cache reuse. Antonio Cardace (@acardace), Principal Machine Learning Engineer at Red Hat, is a dedicated technologist with a deep background in Kubernetes, Virtualization, and embedded development. An active member of the open-source community, he currently contributes to the llm-d project and previously served as a maintainer for the KubeVirt project.
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
this little its bitsy tiny gemma4 model on my 3060 is talking better than chat gpt
Discover how the Gemma4 model outperforms ChatGPT on a 3060 GPU, and why it matters for AI development
Reddit r/artificial
📰
Prompt Injection for API Teams: What It Is and How to Test for It
Learn to identify and test for prompt injection vulnerabilities in API teams to ensure model security
Dev.to · Hassann
📰
How to track new AI drops without the social media delay?
Stay updated on new AI drops without social media delay by following source-of-truth channels and setting up alerts with keywords
Reddit r/artificial
📰
Inside the Model Factory — Eiso Kant, Poolside AI
Learn how Poolside AI's small team built a model factory to train Laguna S, a 118B MOE model that beats Thinky's 1T open weights model
Latent Space
Up next
5 Levels of AI Agents - From Simple LLM Calls to Multi-Agent Systems
Dave Ebbelaar (LLM Eng)
Watch →