30. 09. 2026 Alessandro Taufer Uncategorized

How Kubernetes pods actually end up on a node

Scheduling roulette

You apply a deployment, run kubectl get pods, and there they are: every Pod neatly sitting on some node. It feels a bit like roulette. You spin the wheel and the Pod lands wherever it lands.

The good news is that wheel is not as random as it looks. In this article I’d like to walk you through how the kubernetes scheduler really decides where a Pod goes, and then put it to the test with three small experiments. As always, a couple of concrete examples teach more than any diagram.

Inside the scheduler: three queues

What happens to a Pod that is waiting for a node? It sits in a queue, or rather in one of three:

  • activeQ Pods that are ready to be scheduled, sorted by priority. The scheduler picks them up one at a time.
  • unschedulablePods Pods the scheduler already tried to place and couldn’t. They wait here until something changes in the cluster.
  • backoffQ Pods that are getting a second chance, but only after a short delay, so that they don’t hammer the scheduler.

The happy path is boring: a Pod leaves the activeQ, finds a node and gets bound. The interesting one is the unhappy path. When no node fits, the pod is parked in unschedulablePods until an event comes along that might change the answer, like a new node joining or another Pod being deleted. Then it makes its way back to the activeQ, usually via the backoffQ, for another try. Thanks to a feature called queueing hints, plugins can tell the scheduler which events are actually worth retrying for.
This is also what’s going on behind a Pod stuck in Pending: it hasn’t been forgotten, it’s waiting for the world to change.

Filter first, score later

So what does the scheduler do once it picks a Pod up? Its main job boils down to two steps:

  1. Node selection: throw away every node that cannot run the pod. The ones left standing are called feasible nodes.
  2. Node scoring: rank the survivors according to the active scoring rules and pick the one with the highest score.

If the first step leaves you with zero nodes, the Pod heads to unschedulablePods and we’re back to the queues. If it leaves you with several, scoring takes over.

And what if two nodes end up with the same top score? Here’s where the roulette comes in: the scheduler simply picks one of them at random.

What makes a node feasible?

The selection step looks at a handful of criteria. These are the ones you’ll run into most often:

  • Node resources fit the node must have enough room left for the Pod’s CPU and memory requests. Keep an eye on that last word: we’ll come back to it in the final experiment.
  • Node affinity and anti-affinity nodeSelector, node affinity rules, and (anti-)affinity towards other Pods. Basically “put me on nodes with this label” or “keep me away from Pods like that one”.
  • Taints and tolerations nodes can repel Pods unless the Pod explicitly tolerates the taint. It’s how you keep regular workloads away from special nodes.
  • Topology spread constraints: spread your Pods across failure domains, like nodes or zones, instead of piling them up in one place.
  • Priority and preemption priority decides who goes first in the queue. When a high priority Pod fits nowhere, the scheduler may evict lower priority Pods to make room for it.

Some of these also come with a “soft” flavor: instead of crossing nodes off the list, they just make a node more or less attractive. That’s the difference between you must and you’d better.

Playing three-node monte

Enough theory, let’s play. For the next experiments I’m using a small minikube cluster with three nodes. The rules are simple: I show you a manifest, you guess where the Pods will end up. No peeking at the scheduler logs!

Experiment: The greedy pod

Meet a really bad neighbor:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: noisy-neighbor
spec:
  replicas: 1
  # {...}
  template:
    # {...}
    spec:
      nodeSelector:
        kubernetes.io/hostname: minikube
      containers:
      - name: stress
        image: polinux/stress
        command: ["stress"]
        args:
        - --vm
        - "1"
        - --vm-bytes
        - "1000M"
        - --vm-keep
        resources:
          requests:
            memory: "1Ki" # the lie: the scheduler thinks this Pod costs almost nothing

The stress tool allocates about 1 GB of memory and keeps touching it. The Pod, on the other hand, requests a single kibibyte. And here’s the thing: the scheduler only looks at requests, never at actual usage.

From its point of view, this Pod is practically free. The node still looks empty, so more Pods keep coming while the real memory is long gone. The scheduler did nothing wrong: it trusted the numbers it was given, and the numbers were a lie. (I pinned the Pod to a node to make the demo reproducible, but nothing stops the scheduler from putting a Pod like this right next to your database.)

What’s next?

What if the scheduler could see the real load? That’s the idea behind Trimaran, a collection of load-aware scheduling plugins. Instead of relying on requests alone, they read the actual utilization of your nodes from sources like Metrics Server or Prometheus, through a component called load-watcher. There are three flavors:

  • TargetLoadPacking packs Pods until a target CPU utilization is reached, then switches to spreading.
  • LoadVariationRiskBalancing balances the risk across nodes, a mix of average utilization and how much it fluctuates.
  • LowRiskOverCommitment: takes both the limits and the real load into account, to keep overcommitment under control.

The catch? These plugins live outside the stock kube-scheduler, so you run them as an additional scheduler profile, and they are only as good as the metrics you feed them.

The other topic on my radar is gang scheduling: all-or-nothing placement for groups of Pods that are useless unless they all run, like distributed training jobs. The default scheduler treats every Pod on its own, so half of a job can be running while the other half waits forever. Until recently you needed out-of-tree tools, like the Coscheduling plugin or Volcano. Now Kubernetes has its own answer: the Workload API and a first implementation of gang scheduling arrived as alpha in v1.35, and in v1.37 the Workload and PodGroup APIs graduated to beta, although still behind a feature gate.

The takeaway

The scheduler is not a roulette wheel, and it’s not a fortune teller either. It’s a very diligent accountant: hard constraints cross nodes off the list, scores rank what’s left, ties are settled by chance, and the only thing it trusts is the numbers you wrote in your manifest.

So whenever you write a Pod spec, ask yourself: what am I telling the scheduler, and is it true? If the answer is “not really”, you’ve rigged the game. And sooner or later, one of your nodes pays the price.

Interested in further reading? Have a look at Simplifying Multi-cluster Kubernetes Monitoring with EDOT.

This article was written by AI, based on the “Scheduling Roulette” talk for Let’s Talk IT 2026.

These Solutions are Engineered by Humans

Did you find this article interesting? Does it match your skill set? Programming is at the heart of how we develop customized solutions. In fact, we’re currently hiring for roles just like this and others here at Würth Phoenix.

Author

Alessandro Taufer

Leave a Reply

Your email address will not be published. Required fields are marked *

Archive