When using Kubernetes on bare metal, exposing a LoadBalancer Service usually requires some mechanism to make the Service IP reachable from the surrounding network.
With Cilium, one possible solution is L2 Announcements. Cilium can answer ARP or NDP requests for the LoadBalancer IP, effectively making that IP reachable from the local network.
At first sight, this can create a slightly confusing situation.
Suppose we have the following topology:
Client
|
| 172.19.69.250
v
Node 1 Node 2
|
+-- Application Pod
10.42.2.17
The LoadBalancer IP 172.19.69.250 is currently advertised by Node 1, but the Pod selected by the Service is running on Node 2.
So what actually happens when the client connects to 172.19.69.250?
Does Node 1 somehow proxy the connection? Does the packet first reach a local Kubernetes Service and then get forwarded elsewhere? And what role does VXLAN play in all of this?
The first useful distinction is that advertising a Service IP and load balancing the Service are two different jobs.
Cilium L2 Announcements makes a Service reachable on the local network by having one node answer ARP requests for its IPv4 address, or NDP requests for IPv6.
For example, a machine on the same Layer 2 network might ask:
Who has 172.19.69.250?
If Node 1 is currently responsible for announcing that Service, Cilium answers with the MAC address of one of Node 1’s interfaces.
From the client’s point of view, the result is therefore approximately:
172.19.69.250 -> MAC address of Node 1
The next packet sent to the LoadBalancer IP consequently arrives at Node 1. This is the main purpose of L2 Announcements: getting traffic for the virtual Service IP to one of the Kubernetes nodes.
The IP does not need to be configured as a normal address on the interface. Cilium treats it as a virtual Service IP and responds to the Layer 2 discovery protocol on its behalf.
Cilium’s documentation describes the same model: for a particular Service, one node at a time responds to ARP or NDP requests and then performs Service load balancing for the received traffic.
But reaching Node 1 is only the first half of the journey.
Imagine the client opens a TCP connection to:
172.19.69.250:443
Once the packet reaches Node 1, Cilium recognizes 172.19.69.250:443 as the frontend of a Kubernetes Service.
Conceptually, Cilium has information similar to:
Service frontend
172.19.69.250:443
|
+-- 10.42.1.25:8443
|
+-- 10.42.2.17:8443
|
+-- 10.42.3.31:8443
The exact representation is maintained internally in eBPF maps, but for understanding the packet flow the important part is simply that Cilium knows:
This is where Service load balancing happens.
If the selected backend is:
10.42.2.17:8443
then the Service destination is translated toward that backend.
In a cluster where Cilium is configured with:
kubeProxyReplacement: true
this Service handling is performed by Cilium’s eBPF datapath rather than by kube-proxy installing iptables or IPVS rules. Cilium maintains eBPF maps for Service frontends, backends and related load-balancing state.
So, at this point, we can already make an important distinction:
172.19.69.250:443
answers the question:
Which Service did the client contact?
while:
10.42.2.17:8443
answers:
Which actual workload will handle this connection?
Those are not necessarily located on the same node.
Now we reach the interesting part.
Node 1 received the packet because it currently advertises the LoadBalancer IP, but our selected backend 10.42.2.17 lives on Node 2.
If Cilium is operating in VXLAN encapsulation mode, traffic between nodes is carried through the Cilium overlay.
The path now looks approximately like this:
Kubernetes cluster
Client
|
| dst = 172.19.69.250:443
v
+-----------------------+
| Node 1 |
| |
| L2 announcement |
| brings packet here |
| | |
| v |
| Cilium Service lookup |
| | |
| v |
| backend selected: |
| 10.42.2.17:8443 |
+-----------+-----------+
|
| VXLAN
v
+-----------------------+
| Node 2 |
| |
| 10.42.2.17:8443 |
| | |
| v |
| Pod |
+-----------------------+
Node 1 knows that 10.42.2.17 belongs to another Kubernetes node.
Instead of requiring the physical network to understand routes toward every Pod CIDR, Cilium can encapsulate the packet inside another packet addressed from Node 1 to Node 2.
With VXLAN, the physical network therefore only needs normal node-to-node connectivity. Cilium’s documentation describes encapsulation mode as a mesh of tunnels between cluster nodes, with inter-node traffic transported using VXLAN or Geneve.
The underlay network sees something broadly like:
Node 1 IP -> Node 2 IP
while the encapsulated packet still carries the traffic destined for the selected Pod.
Once Node 2 receives and decapsulates it, Cilium can forward the packet to the Pod.
This separation is quite useful because the external network does not need to know anything about addresses such as 10.42.2.17.
Looking at the entire flow, there are actually three independent questions being answered.
L2 Announcements answers this.
For our example:
172.19.69.250 -> Node 1
Node 1 currently answers ARP for the virtual IP.
Kubernetes Service load balancing answers this.
For example:
172.19.69.250:443
|
v
10.42.2.17:8443
Cilium selects one of the Service backends.
The Cilium routing mode answers this.
In our example, the Pod is remote and the cluster uses VXLAN, so the packet is encapsulated and transported from Node 1 to Node 2.
These mechanisms cooperate, but they solve different problems.
That is why there is no requirement for the node announcing a LoadBalancer IP to also run one of the Service Pods.
In fact, Cilium explicitly notes an important consequence of this design: when using L2 Announcements, all incoming traffic for a given Service initially reaches the node currently advertising the IP. Load distribution happens only after traffic has entered the cluster.
In this particular setup: nowhere.
Traditionally, Kubernetes relies on kube-proxy to implement Service forwarding, commonly through iptables or IPVS.
With Cilium configured for kube-proxy replacement:
kubeProxyReplacement: true
Cilium implements Kubernetes Service handling directly using eBPF, and this is also a prerequisite for Cilium L2 Announcements.
Cilium itself knows that the LoadBalancer address represents a Service and performs the backend lookup and translation.
There is one last piece to consider: the response.
The Pod eventually sends traffic back toward the client. Cilium must make sure that the connection still looks consistent from the client’s perspective.
The exact path depends on the configured Cilium load-balancing mode.
For example, different setups may use SNAT or Direct Server Return, and those choices affect whether reply traffic passes again through the original load-balancing node.
In the common SNAT case, the node performing the load balancing also maintains the translation required for the response path. Cilium’s kube-proxy replacement documentation distinguishes this from DSR, where the backend can reply directly to the client.
That topic deserves an article of its own.
For understanding L2 Announcements, the important point is simply that receiving the original packet and selecting the backend is only part of the Service datapath; Cilium also maintains the state needed to handle the corresponding replies.
Returning to our original example:
Client
|
| ARP: Who has 172.19.69.250?
|
| <--- Node 1 answers
|
| TCP 172.19.69.250:443
v
+--------------------------------------------------+
| Node 1 |
| |
| L2 Announcement |
| | |
| v |
| Cilium eBPF Service load balancing |
| |
| 172.19.69.250:443 -> 10.42.2.17:8443 |
| | |
+---------------------------+----------------------+
|
| VXLAN
|
+---------------------------v----------------------+
| Node 2 |
| |
| Pod |
| 10.42.2.17:8443 |
+--------------------------------------------------+
The node announcing the LoadBalancer IP therefore does not “own” the application.
It owns, temporarily, the responsibility of making that virtual IP reachable from the surrounding Layer 2 network.
After the packet reaches the cluster, Kubernetes Service load balancing decides which backend should receive it. If that backend is on another node, the normal Cilium datapath takes care of reaching it — VXLAN in the example above.
Thinking about these as separate layers makes the whole mechanism much easier to reason about:
L2 Announcement
|
| Gets the packet into the cluster
v
Service Load Balancing
|
| Selects the workload
v
Cilium Routing
|
| Gets the packet to that workload
v
Pod
And that is ultimately the key point:
the node that advertises a Kubernetes LoadBalancer IP and the node that runs the selected workload are two completely separate decisions.