Blog Entries

30. 06. 2026 Davide Sbetti AI, Kubernetes

Load-balancing Requests to LLMs in Kubernetes: A KV-cache Approach with llm-d!

Hi everyone 😃 Today I’d like to walk you through some experiments we ran about load-balancing requests to LLMs in Kubernetes. Let’s dive deep into it! Why Deploy LLMs in Kubernetes? Well, Kubernetes has now established itself as a leading technology when it comes to workloads orchestration. And with time the increasing support for accelerators…

Read More

Archive