Kubernetes networking with Cilium

Cilium kube-proxy replacement, the secure way

You removed kube-proxy, Cilium took over Services in eBPF, and everything got faster. It also quietly answers NodePorts on every address a node has, ignores loadBalancerSourceRanges for anything coming from inside, and keeps your old iptables rules around as a souvenir.

The short answer

With kubeProxyReplacement: true, point k8sServiceHost at the real API server, restrict NodePort and LoadBalancer handling to the private NIC with devices and nodePort.addresses, enable bpf.lbSourceRangeAllTypes so source ranges also guard the NodePort, keep external ClusterIP access off, set socketLB.hostNamespaceOnly for sidecar meshes, and remove every kube-proxy leftover.

Updated Houssam Hammoudi, CTOTested with Cilium 1.20.2, Kubernetes 1.34 (kind, no kube-proxy)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Cilium's kube-proxy replacement implements Services in eBPF: ClusterIP at the socket, NodePort and LoadBalancer on the node's network devices. The defaults aim at "it works everywhere", and a few of them reach further than people expect:

  • NodePorts on every suitable address. By default, NodePort and LoadBalancer services answer on the addresses of devices with the default route or a Kubernetes node IP. On a node with a public interface, that includes the public address.
  • Source ranges only on the LoadBalancer. loadBalancerSourceRanges restricts the LoadBalancer address. The NodePort that Kubernetes creates for the same Service is not restricted unless you say so, so the same backend is reachable around the filter.
  • Nothing filters from inside. Pods and host processes in the cluster can reach a LoadBalancer Service regardless of its source ranges.
  • Leftovers. kube-proxy's iptables rules stay on nodes after it is removed, and a kubeadm upgrade can reinstall kube-proxy if its ConfigMap remains.
  • Sidecar meshes. Socket load balancing inside pods bypasses sidecars; see Cilium with a sidecar mesh.

What the docs say

By default the specified white-listed CIDRs in spec.loadBalancerSourceRanges only apply to the LoadBalancer service, but not the corresponding NodePort or ClusterIP service which get installed along with the LoadBalancer service.

Source: Cilium docs, Kubernetes Without kube-proxy

When accessing the service from inside a cluster, the kube-proxy replacement will ignore the field regardless whether it is set.

Source: Cilium docs, Kubernetes Without kube-proxy

When running Cilium’s eBPF kube-proxy replacement, by default, a NodePort or LoadBalancer service or a service with externalIPs will be accessible through the IP addresses of native devices which have the default route on the host or have Kubernetes InternalIP or ExternalIP assigned.

Source: Cilium docs, Kubernetes Without kube-proxy

Be aware that removing kube-proxy will break existing service connections.

Source: Cilium docs, Kubernetes Without kube-proxy

Everything on this page is in Cilium's docs, spread across a very long page that is mostly about performance. The security-relevant switches are devices, nodePort.addresses, bpf.lbSourceRangeAllTypes and bpf.lbExternalClusterIP, and none of them is in the quick start.

The secure configuration

1. Remove kube-proxy completely (new clusters: never install it).

bash
kubectl -n kube-system delete ds kube-proxy
# Delete the ConfigMap too, so a kubeadm upgrade does not reinstall kube-proxy.
kubectl -n kube-system delete cm kube-proxy
# On each node, as root: drop kube-proxy's iptables rules.
iptables-save | grep -v KUBE | iptables-restore

Removing kube-proxy on a running cluster breaks existing Service connections. Do it in a maintenance window, or build new nodes without it and move workloads.

2. Cilium Helm values.

yaml
kubeProxyReplacement: true
k8sServiceHost: api.k8s.example.com      # the API server itself, not the kubernetes ClusterIP
k8sServicePort: 6443
devices: eth1                            # the private NIC: NodePort and LoadBalancer only here
nodePort:
  addresses:
    - 10.20.0.0/16                       # only node addresses in the private subnet serve NodePorts
bpf:
  lbSourceRangeAllTypes: true            # loadBalancerSourceRanges also guard the NodePort
  lbExternalClusterIP: false             # ClusterIPs stay unreachable from outside the cluster
socketLB:
  hostNamespaceOnly: true                # keep sidecar meshes working (see the mesh page)

These render as devices: "eth1", nodeport-addresses: "10.20.0.0/16", bpf-lb-source-range-all-types: "true", bpf-lb-external-clusterip: "false" and bpf-lb-sock-hostns-only: "true" in the cilium-config ConfigMap. devices must name the same interface on every node, or use a wildcard such as eth+.

3. Services exposed on purpose, with no side door. For a LoadBalancer that should only be reachable from known ranges, stop the companion NodePort and ClusterIP frontends from being created:

yaml
apiVersion: v1
kind: Service
metadata:
  name: admin-api
  namespace: app
  annotations:
    service.cilium.io/type: LoadBalancer     # no side-door NodePort or ClusterIP for this Service
spec:
  type: LoadBalancer
  allocateLoadBalancerNodePorts: false       # Kubernetes allocates no node port at all
  loadBalancerSourceRanges:
    - 203.0.113.0/24                         # office egress range
  selector:
    app.kubernetes.io/name: admin-api
  ports:
    - name: https
      port: 443
      targetPort: 8443

Source ranges do not apply to traffic from inside the cluster. Protect the backend pods with a network policy as well.

Prove it

The lab is a kind cluster without kube-proxy. Each node got a second NIC, eth1 on 192.0.2.0/24, and the default route moved to it, the way a cloud VM has a public NIC with the default route. In this lab the private NIC is eth0 on 172.18.0.0/16, so the values below use devices: eth0 and nodePort.addresses: [172.18.0.0/16] where your nodes use eth1 and 10.20.0.0/16.

1. Before: auto-detection publishes NodePorts on the public NIC.

text
$ cilium-dbg status --verbose | grep -A14 'KubeProxyReplacement Details' | grep Devices
  Devices:              eth0    172.18.0.2 ... (Direct Routing), eth1  192.0.2.11
$ curl -s -o /dev/null -w '%{http_code}\n' http://172.18.0.2:30080/
200
$ curl -s -o /dev/null -w '%{http_code}\n' http://192.0.2.11:30080/
200

Nobody configured eth1. It was picked because it holds the default route.

2. After: the replacement runs on the private device only.

text
$ kubectl -n kube-system get cm cilium-config -o json | jq -r '.data | ...'
bpf-lb-external-clusterip: false
bpf-lb-sock-hostns-only: true
bpf-lb-source-range-all-types: true
devices: eth0
nodeport-addresses: 172.18.0.0/16
$ cilium-dbg status --verbose | grep -A14 'KubeProxyReplacement Details' | grep -E 'Status|Coverage|Devices'
  Status:               True
  Socket LB Coverage:   Hostns-only
  Devices:              eth0    172.18.0.2 ... (Direct Routing)
$ curl -s -o /dev/null -m 4 -w '%{http_code}\n' http://172.18.0.2:30080/
200
$ curl -s -o /dev/null -m 4 -w '%{http_code}\n' http://192.0.2.11:30080/
000

000 is curl giving up after 4 seconds: nothing answers on the public address. cilium-dbg service list does not help here. In Cilium 1.20 it shows one 0.0.0.0:30080/TCP NodePort frontend whatever the devices are, so test with a real connection from outside.

nodePort.addresses is a second, separate layer. With devices: eth+ (both NICs back in), the public address still did not answer, because it is outside 172.18.0.0/16.

3. The source range also guards the NodePort. A LoadBalancer Service with loadBalancerSourceRanges: [203.0.113.0/24] and node port 30081, tested from a host outside that range:

text
bpf.lbSourceRangeAllTypes=false:  curl http://172.18.0.2:30081/  ->  200
bpf.lbSourceRangeAllTypes=true:   curl http://172.18.0.2:30081/  ->  000 (timeout)
add the host to the ranges:       curl http://172.18.0.2:30081/  ->  200

With the default, the filter on the LoadBalancer address is decoration: the node port next to it serves anyone.

4. The LoadBalancer-only Service has no side doors.

text
$ kubectl -n edge-test get svc admin-api -o jsonpath='{.spec.type} nodePort=[{.spec.ports[0].nodePort}]'
LoadBalancer nodePort=[]
$ cilium-dbg service list | grep -E '10.96.6.12|172.18.250'
31   172.18.250.0:80/TCP      LoadBalancer   1 => 10.244.2.159:80/TCP (active)
32   172.18.250.1:80/TCP      LoadBalancer   1 => 10.244.2.159:80/TCP (active)

10.96.6.12 is the ClusterIP of admin-api, and it has no frontend. From a pod in the cluster:

text
curl http://admin-api.edge-test.svc.cluster.local/   ->  000 (timeout)
curl http://ranged.edge-test.svc.cluster.local/      ->  200
curl http://10.244.2.159/                            ->  200

The last line is the backend pod, reached directly. The annotation removes the Service's doors, not the pod's. That is why the page asks for network policy on the backend pods.

5. No kube-proxy leftovers:

text
$ kubectl -n kube-system get ds,cm kube-proxy
Error from server (NotFound): daemonsets.apps "kube-proxy" not found
Error from server (NotFound): configmaps "kube-proxy" not found
$ for n in lab-control-plane lab-worker lab-worker2; do docker exec $n sh -c "iptables-save | grep -c KUBE-SVC"; done
0
0
0

Mistakes people make

Letting Cilium pick the devices

Auto-detection picks devices with the default route, which on many nodes is the public one. Name the private device.

Trusting loadBalancerSourceRanges alone

It guards the LoadBalancer address only, by default, and never guards traffic from inside the cluster. Enable bpf.lbSourceRangeAllTypes, use the service.cilium.io/type annotation for sensitive Services, and add network policy.

Enabling external ClusterIP access for convenience

bpf.lbExternalClusterIP: true makes every ClusterIP reachable from outside the cluster through the nodes. Services that were never meant to be exposed suddenly are.

Changing one Helm value and the chart version with it

helm upgrade --reuse-values across chart versions can fail on the schema (the lab got at '/endpointPolicyUpdateTimeoutDuration': got string, want null). Use --reset-then-reuse-values. And pass the --version that is running now (helm list -n kube-system): an older one is a silent downgrade. In the lab that took two agents into CrashLoopBackOff with Strict ingress encryption requires tunneling to be enabled., a setting the newer version accepts.

Leaving kube-proxy's ConfigMap

A kubeadm upgrade can bring kube-proxy back next to Cilium, with two NAT systems that do not know about each other. Delete the ConfigMap.

Pointing k8sServiceHost at the kubernetes Service

Without kube-proxy, nothing serves that ClusterIP until Cilium is up. Point Cilium at the API server's real address or a stable DNS name.

Checklist

  • kube-proxy's DaemonSet and ConfigMap are deleted, and no KUBE- iptables rules remain.
  • k8sServiceHost points at the API server's real address or DNS name.
  • devices names the private interface only.
  • nodePort.addresses lists only private node subnets.
  • bpf.lbSourceRangeAllTypes is true.
  • bpf.lbExternalClusterIP is false.
  • socketLB.hostNamespaceOnly is true if any sidecar mesh runs.
  • Sensitive LoadBalancer Services use service.cilium.io/type: LoadBalancer and source ranges.
  • Backend pods of exposed Services also have network policy.
  • NodePorts do not answer on public node addresses.

Replacing kube-proxy is a performance win you should take. Just make sure the eBPF version publishes exactly what the iptables version did, and not the extra doors it found along the way.

H2-CSPE

Learn it on a live range

Cluster networking and policy, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on kubernetes networking with cilium

Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.

All kubernetes networking with cilium guides