Kubernetes networking with Cilium

Cilium with minimal capabilities on Talos, the secure way

Talos took away SSH, the shell and the package manager. Then the default Cilium chart arrived with a privileged init container, two more that nsenter into the host, and a capability to load kernel modules on a system that does not let workloads load kernel modules.

The short answer

On Talos, install Cilium with SYS_MODULE and SYSLOG removed from the agent's capabilities, cgroup and bpffs auto-mount disabled because Talos already mounts them, the sysctl fix init container off, and KubePrism on localhost:7445 as the API endpoint. No Cilium container is privileged. Then deny pods/exec on the agent: exec there is root on the node.

Updated Houssam Hammoudi, CTOTested with Cilium 1.20.2 on Kubernetes 1.34 (kind, nodes with bpffs and cgroup2 already mounted); Talos docs (Sidero)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Cilium needs real privileges: it programs eBPF, routes and network namespaces on the node. The chart's defaults are written for every Linux distribution, so they include privileges for work that Talos already does:

  • SYS_MODULE on the agent and the state-cleaning init container, to load kernel modules. Talos does not allow workloads to load modules.
  • A privileged mount-bpf-fs init container to mount the BPF filesystem, which Talos already mounts.
  • mount-cgroup and apply-sysctl-overwrites init containers with SYS_ADMIN, SYS_CHROOT and SYS_PTRACE, which mount the host's /proc and use nsenter to act in the host's namespaces.
  • SYSLOG to read the kernel log.

Every extra privilege is something an attacker gets if they gain code execution in that container, through a Cilium bug or through kubectl exec by someone who should not have it. On Talos, where the host itself has no shell, the Cilium agent is one of the few ways left to act as root on a node.

What the docs say

Talos does not allow Kubernetes workloads to load kernel modules, so the SYS_MODULE capability must be dropped from Cilium’s default capability set

Source: Sidero docs, Deploy Cilium CNI

Talos already provides cgroupv2 and bpffs mounts — do not let Cilium attempt to mount these

Source: Sidero docs, Deploy Cilium CNI

If pod exec operations aren’t restricted, then remote exec into pods and containers defeats Linux namespace restrictions.

Source: Cilium docs, Restricting privileged Cilium pod access

To prevent privileged access to Cilium pods, restrict access to the Kubernetes API and arbitrary pod exec operations.

Source: Cilium docs, Restricting privileged Cilium pod access

The Talos guide's example values turn off the cgroup mount but not the BPF filesystem mount, even though its own prerequisites say Talos provides both. Turning the BPF mount off too is right, but only together with the Envoy mount on this page. Neither guide mentions the sysctl init container, and Cilium's page on exec access stops at two lines: configure RBAC, and limit access to the nodes proxy subresource.

The secure configuration

1. Talos machine config: no default CNI, no kube-proxy, KubePrism on. Follow the Talos guide for your Talos version (the patch format changed in v1.14); KubePrism listens on localhost:7445 on every node, so Cilium never depends on a ClusterIP to reach the API server.

2. Cilium Helm values for Talos.

yaml
ipam:
  mode: kubernetes
kubeProxyReplacement: true
k8sServiceHost: localhost          # KubePrism on every Talos node
k8sServicePort: 7445
securityContext:
  privileged: false
  capabilities:
    ciliumAgent:                   # Cilium's default list without SYS_MODULE and SYSLOG
      - CHOWN
      - KILL
      - NET_ADMIN
      - NET_RAW
      - IPC_LOCK
      - SYS_ADMIN
      - SYS_RESOURCE
      - DAC_OVERRIDE
      - FOWNER
      - SETGID
      - SETUID
    cleanCiliumState:
      - NET_ADMIN
      - SYS_ADMIN
      - SYS_RESOURCE
bpf:
  autoMount:
    enabled: false                 # Talos mounts bpffs: no privileged mount-bpf-fs init container
envoy:
  extraHostPathMounts:             # bpf.autoMount.enabled=false also removes this mount from cilium-envoy
    - name: bpf-maps
      mountPath: /sys/fs/bpf
      hostPath: /sys/fs/bpf
      hostPathType: Directory
      mountPropagation: HostToContainer
cgroup:
  autoMount:
    enabled: false                 # Talos mounts cgroup2: no mount-cgroup init container (nsenter, SYS_CHROOT, SYS_PTRACE)
  hostRoot: /sys/fs/cgroup
sysctlfix:
  enabled: false                   # no apply-sysctl-overwrites init container (it mounts the host /proc to edit /etc/sysctl.d); Talos sets sysctls in its machine config

The envoy block is not optional. In the 1.20.2 chart, the switch that removes the privileged init container also removes the BPF filesystem mount from the cilium-envoy DaemonSet. Envoy then cannot open Cilium's maps, and every connection that an HTTP-aware policy sends to it is reset (see Prove it, step 3).

SYS_ADMIN stays. The chart's comments suggest BPF and PERFMON could replace it on recent kernels, but the same comments say it is still needed to switch network namespaces. Do not remove it without testing every feature you use.

3. No exec into Cilium pods. Give operators a read-only role in kube-system, and make sure nothing broader grants pods/exec, pods/attach or nodes/proxy:

yaml
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: cilium-readonly
  namespace: kube-system
rules:
  - apiGroups: [""]
    resources: ["pods", "pods/log"]
    verbs: ["get", "list", "watch"]

For troubleshooting that needs cilium-dbg, use a break-glass group with an expiry and an audit trail.

Prove it

1. What the chart renders, before and after. This ran in a container: helm template of the Cilium 1.20.2 chart, listing each container of the agent DaemonSet (see secure-tests/). With the default security settings:

text
config                 privileged=False drop=['ALL'] add=['NET_ADMIN']
mount-cgroup           privileged=False drop=['ALL'] add=['SYS_ADMIN', 'SYS_CHROOT', 'SYS_PTRACE']
apply-sysctl-overwrites privileged=False drop=['ALL'] add=['SYS_ADMIN', 'SYS_CHROOT', 'SYS_PTRACE']
mount-bpf-fs           privileged=True  drop=None add=None
clean-cilium-state     privileged=False drop=['ALL'] add=['NET_ADMIN', 'SYS_MODULE', 'SYS_ADMIN', 'SYS_RESOURCE']
install-cni-binaries   privileged=False drop=['ALL'] add=None
cilium-agent           privileged=False drop=['ALL'] add=['CHOWN', 'KILL', 'NET_ADMIN', 'NET_RAW', 'IPC_LOCK', 'SYS_MODULE', 'SYS_ADMIN', 'SYS_RESOURCE', 'DAC_OVERRIDE', 'FOWNER', 'SETGID', 'SETUID', 'SYSLOG']

With the Talos values above:

text
config                 privileged=False drop=['ALL'] add=['NET_ADMIN']
clean-cilium-state     privileged=False drop=['ALL'] add=['NET_ADMIN', 'SYS_ADMIN', 'SYS_RESOURCE']
install-cni-binaries   privileged=False drop=['ALL'] add=None
cilium-agent           privileged=False drop=['ALL'] add=['CHOWN', 'KILL', 'NET_ADMIN', 'NET_RAW', 'IPC_LOCK', 'SYS_ADMIN', 'SYS_RESOURCE', 'DAC_OVERRIDE', 'FOWNER', 'SETGID', 'SETUID']

Three init containers are gone, including the only privileged one, and SYS_MODULE and SYSLOG are gone from the agent.

Steps 2 and 3 ran on a lab cluster that is not Talos: kind nodes, which also have bpf and cgroup2 mounted before Cilium starts. The values were applied with helm upgrade to a cluster already using WireGuard, network policies, FQDN rules, L2 announcements and an HTTP-aware policy, with k8sServiceHost left at the lab's API server instead of KubePrism.

2. The running pods match the render, and the agent is healthy:

text
$ kubectl -n kube-system get pod -l k8s-app=cilium -o json | jq -c '[.items[0].spec.initContainers[], .items[0].spec.containers[]] | map({name, priv: .securityContext.privileged, caps: .securityContext.capabilities.add})[]'
{"name":"config","priv":null,"caps":["NET_ADMIN"]}
{"name":"clean-cilium-state","priv":null,"caps":["NET_ADMIN","SYS_ADMIN","SYS_RESOURCE"]}
{"name":"install-cni-binaries","priv":null,"caps":null}
{"name":"cilium-agent","priv":null,"caps":["CHOWN","KILL","NET_ADMIN","NET_RAW","IPC_LOCK","SYS_ADMIN","SYS_RESOURCE","DAC_OVERRIDE","FOWNER","SETGID","SETUID"]}
$ cilium-dbg status
Cilium:                  Ok   1.20.2 (v1.20.2-e0dc92bd)
Controller Status:       141/141 healthy
Encryption:              Wireguard       [NodeEncryption: Enabled, cilium_wg0 (..., Peers: 2)]

Policies, FQDN rules and L2 announcements kept working. New pod interfaces kept rp_filter = 0 without the sysctl init container. That container exists because systemd 245 and later set rp_filter to 1 on every interface (Cilium PR #20072), and Talos does not run systemd.

3. HTTP-aware policies need the Envoy mount. A policy that allows only GET / to a web pod, tested from another pod:

text
bpf.autoMount false, no envoy block:   GET /  -> 000 (curl exit 56, connection reset)
                                       GET /admin -> 000 (curl exit 56)
bpf.autoMount false, envoy block:      GET /  -> 200
                                       GET /admin -> 403

The broken state is quiet: the pods are Running, and only the Envoy log says what is wrong:

text
[warning][filter] [cilium/conntrack.cc:86] cilium.bpf_metadata: Cannot open IPv4 conntrack map at /sys/fs/bpf/tc/globals/cilium_ct4_global
[warning][filter] [cilium/ipcache.cc:113] cilium.ipcache: Cannot open ipcache at /sys/fs/bpf/tc/globals/cilium_ipcache_v2

The chart template shows why: the Envoy DaemonSet mounts /sys/fs/bpf only inside {{- if .Values.bpf.autoMount.enabled }} (templates/cilium-envoy/daemonset.yaml).

On Talos itself, also confirm the mounts the values rely on:

bash
talosctl -n 192.0.2.21 read /proc/mounts | grep -E ' /sys/fs/(bpf|cgroup) '

4. Operators cannot exec into the agent:

bash
kubectl auth can-i create pods/exec -n kube-system [email protected]
kubectl auth can-i get nodes/proxy [email protected]

What you should see: no for both.

Mistakes people make

Keeping SYS_MODULE "because the default has it"

On Talos it does nothing useful, and anywhere else it lets a compromised agent load kernel code. Drop it.

Turning bpf.autoMount off and forgetting Envoy

Talos mounts bpffs itself, so bpf.autoMount.enabled: false removes a privileged init container for nothing lost. But the same switch drops the bpffs mount from cilium-envoy, and every HTTP-aware policy starts resetting connections. Add the envoy.extraHostPathMounts entry, and test one L7 rule after the change.

Trusting allowPrivilegeEscalation in the values

The chart's values.yaml has securityContext.allowPrivilegeEscalation: false, commented "disable privilege escalation". No template in the 1.20.2 chart reads it, and the agent's process shows NoNewPrivs: 0. It could not work anyway:

allowPrivilegeEscalation: false is inconsistent, and therefore cannot be set in combination, with a container that: is run as privileged, or has the capability CAP_SYS_ADMIN

Source: Kubernetes docs, Configure a Security Context

The agent keeps SYS_ADMIN. Count on the capability list, not on this flag.

Removing SYS_ADMIN to reach "minimal"

It is still required for namespace operations. A broken agent is not more secure, it is just broken. Remove only what you can prove unused.

Letting everyone exec into kube-system

Exec into the Cilium agent is root on the node. Cilium's own guide says to restrict it. Check aggregated roles and CI service accounts, not only humans.

Pointing Cilium at the kubernetes ClusterIP

Without kube-proxy, that address does not work until Cilium runs. On Talos, KubePrism on localhost:7445 is the reliable endpoint.

Checklist

  • The agent's capabilities exclude SYS_MODULE and SYSLOG.
  • cleanCiliumState capabilities exclude SYS_MODULE.
  • bpf.autoMount.enabled and cgroup.autoMount.enabled are false, with cgroup.hostRoot: /sys/fs/cgroup.
  • cilium-envoy has /sys/fs/bpf through envoy.extraHostPathMounts, and one HTTP-aware policy was tested after the change.
  • sysctlfix.enabled is false.
  • No container in the Cilium DaemonSet is privileged.
  • k8sServiceHost: localhost and k8sServicePort: 7445 (KubePrism).
  • Only a break-glass group can exec into Cilium pods or use nodes/proxy.
  • cilium-dbg status is healthy after the change.

Talos removed the easy ways to become root on a node. Trimming Cilium to what it really uses removes one of the few that were left.

H2-CSPE

Learn it on a live range

Cluster networking and policy, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on kubernetes networking with cilium

Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.

All kubernetes networking with cilium guides