Kubernetes networking with Cilium
Cilium with minimal capabilities on Talos, the secure way
Talos took away SSH, the shell and the package manager. Then the default Cilium chart arrived with a privileged init container, two more that nsenter into the host, and a capability to load kernel modules on a system that does not let workloads load kernel modules.
The short answer
On Talos, install Cilium with SYS_MODULE and SYSLOG removed from the agent's capabilities, cgroup and bpffs auto-mount disabled because Talos already mounts them, the sysctl fix init container off, and KubePrism on localhost:7445 as the API endpoint. No Cilium container is privileged. Then deny pods/exec on the agent: exec there is root on the node.
On this page
What goes wrong
Cilium needs real privileges: it programs eBPF, routes and network namespaces on the node. The chart's defaults are written for every Linux distribution, so they include privileges for work that Talos already does:
SYS_MODULEon the agent and the state-cleaning init container, to load kernel modules. Talos does not allow workloads to load modules.- A privileged
mount-bpf-fsinit container to mount the BPF filesystem, which Talos already mounts. mount-cgroupandapply-sysctl-overwritesinit containers withSYS_ADMIN,SYS_CHROOTandSYS_PTRACE, which mount the host's/procand usensenterto act in the host's namespaces.SYSLOGto read the kernel log.
Every extra privilege is something an attacker gets if they gain code
execution in that container, through a Cilium bug or through kubectl exec
by someone who should not have it. On Talos, where the host itself has no
shell, the Cilium agent is one of the few ways left to act as root on a node.
What the docs say
Talos does not allow Kubernetes workloads to load kernel modules, so the SYS_MODULE capability must be dropped from Cilium’s default capability set
Source: Sidero docs, Deploy Cilium CNI
Talos already provides cgroupv2 and bpffs mounts — do not let Cilium attempt to mount these
Source: Sidero docs, Deploy Cilium CNI
If pod exec operations aren’t restricted, then remote exec into pods and containers defeats Linux namespace restrictions.
Source: Cilium docs, Restricting privileged Cilium pod access
To prevent privileged access to Cilium pods, restrict access to the Kubernetes API and arbitrary pod exec operations.
Source: Cilium docs, Restricting privileged Cilium pod access
The Talos guide's example values turn off the cgroup mount but not the BPF filesystem mount, even though its own prerequisites say Talos provides both. Turning the BPF mount off too is right, but only together with the Envoy mount on this page. Neither guide mentions the sysctl init container, and Cilium's page on exec access stops at two lines: configure RBAC, and limit access to the nodes proxy subresource.
The secure configuration
1. Talos machine config: no default CNI, no kube-proxy, KubePrism on.
Follow the Talos guide for your Talos version (the patch format changed in
v1.14); KubePrism listens on localhost:7445 on every node, so Cilium never
depends on a ClusterIP to reach the API server.
2. Cilium Helm values for Talos.
ipam:
mode: kubernetes
kubeProxyReplacement: true
k8sServiceHost: localhost # KubePrism on every Talos node
k8sServicePort: 7445
securityContext:
privileged: false
capabilities:
ciliumAgent: # Cilium's default list without SYS_MODULE and SYSLOG
- CHOWN
- KILL
- NET_ADMIN
- NET_RAW
- IPC_LOCK
- SYS_ADMIN
- SYS_RESOURCE
- DAC_OVERRIDE
- FOWNER
- SETGID
- SETUID
cleanCiliumState:
- NET_ADMIN
- SYS_ADMIN
- SYS_RESOURCE
bpf:
autoMount:
enabled: false # Talos mounts bpffs: no privileged mount-bpf-fs init container
envoy:
extraHostPathMounts: # bpf.autoMount.enabled=false also removes this mount from cilium-envoy
- name: bpf-maps
mountPath: /sys/fs/bpf
hostPath: /sys/fs/bpf
hostPathType: Directory
mountPropagation: HostToContainer
cgroup:
autoMount:
enabled: false # Talos mounts cgroup2: no mount-cgroup init container (nsenter, SYS_CHROOT, SYS_PTRACE)
hostRoot: /sys/fs/cgroup
sysctlfix:
enabled: false # no apply-sysctl-overwrites init container (it mounts the host /proc to edit /etc/sysctl.d); Talos sets sysctls in its machine configThe envoy block is not optional. In the 1.20.2 chart, the switch that
removes the privileged init container also removes the BPF filesystem mount
from the cilium-envoy DaemonSet. Envoy then cannot open Cilium's maps, and
every connection that an HTTP-aware policy sends to it is reset (see Prove
it, step 3).
SYS_ADMIN stays. The chart's comments suggest BPF and PERFMON could
replace it on recent kernels, but the same comments say it is still needed
to switch network namespaces. Do not remove it without testing every feature
you use.
3. No exec into Cilium pods. Give operators a read-only role in
kube-system, and make sure nothing broader grants pods/exec,
pods/attach or nodes/proxy:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: cilium-readonly
namespace: kube-system
rules:
- apiGroups: [""]
resources: ["pods", "pods/log"]
verbs: ["get", "list", "watch"]For troubleshooting that needs cilium-dbg, use a break-glass group with an
expiry and an audit trail.
Prove it
1. What the chart renders, before and after. This ran in a container:
helm template of the Cilium 1.20.2 chart, listing each container of the
agent DaemonSet (see secure-tests/). With the default security settings:
config privileged=False drop=['ALL'] add=['NET_ADMIN']
mount-cgroup privileged=False drop=['ALL'] add=['SYS_ADMIN', 'SYS_CHROOT', 'SYS_PTRACE']
apply-sysctl-overwrites privileged=False drop=['ALL'] add=['SYS_ADMIN', 'SYS_CHROOT', 'SYS_PTRACE']
mount-bpf-fs privileged=True drop=None add=None
clean-cilium-state privileged=False drop=['ALL'] add=['NET_ADMIN', 'SYS_MODULE', 'SYS_ADMIN', 'SYS_RESOURCE']
install-cni-binaries privileged=False drop=['ALL'] add=None
cilium-agent privileged=False drop=['ALL'] add=['CHOWN', 'KILL', 'NET_ADMIN', 'NET_RAW', 'IPC_LOCK', 'SYS_MODULE', 'SYS_ADMIN', 'SYS_RESOURCE', 'DAC_OVERRIDE', 'FOWNER', 'SETGID', 'SETUID', 'SYSLOG']With the Talos values above:
config privileged=False drop=['ALL'] add=['NET_ADMIN']
clean-cilium-state privileged=False drop=['ALL'] add=['NET_ADMIN', 'SYS_ADMIN', 'SYS_RESOURCE']
install-cni-binaries privileged=False drop=['ALL'] add=None
cilium-agent privileged=False drop=['ALL'] add=['CHOWN', 'KILL', 'NET_ADMIN', 'NET_RAW', 'IPC_LOCK', 'SYS_ADMIN', 'SYS_RESOURCE', 'DAC_OVERRIDE', 'FOWNER', 'SETGID', 'SETUID']Three init containers are gone, including the only privileged one, and
SYS_MODULE and SYSLOG are gone from the agent.
Steps 2 and 3 ran on a lab cluster that is not Talos: kind nodes, which
also have bpf and cgroup2 mounted before Cilium starts. The values were
applied with helm upgrade to a cluster already using WireGuard, network
policies, FQDN rules, L2 announcements and an HTTP-aware policy, with
k8sServiceHost left at the lab's API server instead of KubePrism.
2. The running pods match the render, and the agent is healthy:
$ kubectl -n kube-system get pod -l k8s-app=cilium -o json | jq -c '[.items[0].spec.initContainers[], .items[0].spec.containers[]] | map({name, priv: .securityContext.privileged, caps: .securityContext.capabilities.add})[]'
{"name":"config","priv":null,"caps":["NET_ADMIN"]}
{"name":"clean-cilium-state","priv":null,"caps":["NET_ADMIN","SYS_ADMIN","SYS_RESOURCE"]}
{"name":"install-cni-binaries","priv":null,"caps":null}
{"name":"cilium-agent","priv":null,"caps":["CHOWN","KILL","NET_ADMIN","NET_RAW","IPC_LOCK","SYS_ADMIN","SYS_RESOURCE","DAC_OVERRIDE","FOWNER","SETGID","SETUID"]}
$ cilium-dbg status
Cilium: Ok 1.20.2 (v1.20.2-e0dc92bd)
Controller Status: 141/141 healthy
Encryption: Wireguard [NodeEncryption: Enabled, cilium_wg0 (..., Peers: 2)]Policies, FQDN rules and L2 announcements kept working. New pod interfaces
kept rp_filter = 0 without the sysctl init container. That container
exists because systemd 245 and later set rp_filter to 1 on every interface
(Cilium PR #20072), and Talos does not run systemd.
3. HTTP-aware policies need the Envoy mount. A policy that allows only
GET / to a web pod, tested from another pod:
bpf.autoMount false, no envoy block: GET / -> 000 (curl exit 56, connection reset)
GET /admin -> 000 (curl exit 56)
bpf.autoMount false, envoy block: GET / -> 200
GET /admin -> 403The broken state is quiet: the pods are Running, and only the Envoy log
says what is wrong:
[warning][filter] [cilium/conntrack.cc:86] cilium.bpf_metadata: Cannot open IPv4 conntrack map at /sys/fs/bpf/tc/globals/cilium_ct4_global
[warning][filter] [cilium/ipcache.cc:113] cilium.ipcache: Cannot open ipcache at /sys/fs/bpf/tc/globals/cilium_ipcache_v2The chart template shows why: the Envoy DaemonSet mounts /sys/fs/bpf only
inside {{- if .Values.bpf.autoMount.enabled }}
(templates/cilium-envoy/daemonset.yaml).
On Talos itself, also confirm the mounts the values rely on:
talosctl -n 192.0.2.21 read /proc/mounts | grep -E ' /sys/fs/(bpf|cgroup) '4. Operators cannot exec into the agent:
kubectl auth can-i create pods/exec -n kube-system [email protected]
kubectl auth can-i get nodes/proxy [email protected]What you should see: no for both.
Mistakes people make
Keeping SYS_MODULE "because the default has it"
On Talos it does nothing useful, and anywhere else it lets a compromised agent load kernel code. Drop it.
Turning bpf.autoMount off and forgetting Envoy
Talos mounts bpffs itself, so bpf.autoMount.enabled: false removes a
privileged init container for nothing lost. But the same switch drops the
bpffs mount from cilium-envoy, and every HTTP-aware policy starts resetting
connections. Add the envoy.extraHostPathMounts entry, and test one L7 rule
after the change.
Trusting allowPrivilegeEscalation in the values
The chart's values.yaml has securityContext.allowPrivilegeEscalation: false,
commented "disable privilege escalation". No template in the 1.20.2 chart
reads it, and the agent's process shows NoNewPrivs: 0. It could not work
anyway:
allowPrivilegeEscalation: falseis inconsistent, and therefore cannot be set in combination, with a container that: is run as privileged, or has the capabilityCAP_SYS_ADMIN
Source: Kubernetes docs, Configure a Security Context
The agent keeps SYS_ADMIN. Count on the capability list, not on this flag.
Removing SYS_ADMIN to reach "minimal"
It is still required for namespace operations. A broken agent is not more secure, it is just broken. Remove only what you can prove unused.
Letting everyone exec into kube-system
Exec into the Cilium agent is root on the node. Cilium's own guide says to restrict it. Check aggregated roles and CI service accounts, not only humans.
Pointing Cilium at the kubernetes ClusterIP
Without kube-proxy, that address does not work until Cilium runs. On Talos,
KubePrism on localhost:7445 is the reliable endpoint.
Checklist
- The agent's capabilities exclude
SYS_MODULEandSYSLOG. cleanCiliumStatecapabilities excludeSYS_MODULE.bpf.autoMount.enabledandcgroup.autoMount.enabledarefalse, withcgroup.hostRoot: /sys/fs/cgroup.cilium-envoyhas/sys/fs/bpfthroughenvoy.extraHostPathMounts, and one HTTP-aware policy was tested after the change.sysctlfix.enabledisfalse.- No container in the Cilium DaemonSet is privileged.
k8sServiceHost: localhostandk8sServicePort: 7445(KubePrism).- Only a break-glass group can exec into Cilium pods or use
nodes/proxy. cilium-dbg statusis healthy after the change.
Talos removed the easy ways to become root on a node. Trimming Cilium to what it really uses removes one of the few that were left.
H2-CSPE
Learn it on a live range
Cluster networking and policy, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on kubernetes networking with cilium
Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.
All kubernetes networking with cilium guides