Cilium with a sidecar mesh (Kuma, Istio, Linkerd) without losing mTLS, the secure way
We know why you're here. Kuma swears your traffic is mTLS, the dashboard is green, and tcpdump just showed you a password in plain text. You're not crazy, you heard right: Cilium did it, very politely, before your sidecar ever saw the packet.
The short answer
When Cilium replaces kube-proxy, set socketLB.hostNamespaceOnly=true so connections inside pods keep the Service ClusterIP and the sidecar can match it. Set cni.exclusive=false for the mesh CNI. Run the mesh in strict mTLS and turn off plaintext passthrough so a bypass fails closed. Then prove encryption on the wire with tcpdump and proxy stats.
On this page
What goes wrong
Cilium's kube-proxy replacement load-balances Services at the socket. A BPF
program attached to the cgroup runs on connect() (and sendmsg() for UDP).
When an application connects to a ClusterIP, the program picks a backend and
rewrites the destination to that backend's pod IP. This happens before the
first packet exists.
A sidecar mesh works one step later. Inside the pod, iptables rules (from an init container or the mesh CNI) redirect outbound TCP to the sidecar. The sidecar reads the original destination and matches it against the Services it knows, by ClusterIP and port. That match selects the upstream, and the upstream carries the mTLS settings.
After the socket rewrite, the original destination is a pod IP, not the ClusterIP. The sidecar has no listener for that address. Kuma documents what happens next: an unknown destination goes through passthrough, and passthrough is not encrypted. The connection leaves the pod as plain TCP. Telemetry, retries, timeouts and any policy that needs the client's certificate never see it either.
What happens at the server depends on its mTLS mode:
- Strict: the server sidecar expects TLS, receives plaintext, and closes the connection. The client gets an empty reply. But the request, with any password or token in it, has already crossed the network in clear. The usual "fix" for the errors is to switch to permissive.
- Permissive: the server sidecar accepts plaintext. Everything works, the mesh dashboard stays green, and anyone who can capture on a node, a hypervisor or the network between nodes can read the traffic.
Linkerd is the exception worth knowing. Its proxy also resolves pod IPs, so Linkerd's docs say mTLS and telemetry keep working. What you lose there is Linkerd's load balancing and dynamic request routing.
Nothing logs an error. The trap is on whenever socket load balancing runs in
pod namespaces: with kubeProxyReplacement: true, or with
socketLB.enabled: true next to kube-proxy.
What the docs say
These settings prevent Cilium’s socket-based load balancing from interfering with Istio’s proxying.
Source: Cilium docs, Integration with Istio
Without this option, when Cilium does service resolution via socket load balancing, Istio sidecar will be bypassed, resulting in loss of Istio features including encryption and telemetry.
Source: Cilium 1.13 docs, Getting Started Using Istio
If you pass requests to direct IP addresses, Envoy considers them unknown destinations and manages them in passthrough mode – which means they’re not encrypted with mTLS.
Source: Kuma docs, Kubernetes annotations and labels (kuma.io/direct-access-services)
Consequentially, while mTLS and telemetry will still function correctly, features such as peak EWMA load balancing, and dynamic request routing may not work as expected.
Source: Linkerd docs, Cluster Configuration (Cilium)
Cilium names the setting for Istio only, and the sentence that says
"encryption" was dropped from the current page; it survives in the 1.13 docs.
Kuma's transparent proxying page
does not mention Cilium at all, and Kuma's
CNI page only tells
Cilium users to set cni.exclusive to false. Kuma explains the plaintext
passthrough under an unrelated annotation. No page shows how to check the wire.
The secure configuration
1. Cilium: keep socket load balancing in the host namespace.
# cilium-values.yaml (Helm)
kubeProxyReplacement: true
k8sServiceHost: 192.0.2.10 # your API server address
k8sServicePort: 6443
socketLB:
# Socket LB stays in the host namespace. Inside pods, connect() keeps the
# ClusterIP, so the sidecar can match the Service. Pod traffic that is not
# meshed is still load-balanced by Cilium's tc program on the veth.
# Renders as bpf-lb-sock-hostns-only: "true" in the cilium-config ConfigMap.
hostNamespaceOnly: true
cni:
# Do not delete other CNI configs: kuma-cni, istio-cni and linkerd-cni
# chain after Cilium. Renders as cni-exclusive: "false".
exclusive: false# Use helm, not "cilium upgrade --reuse-values" (see Mistakes).
helm upgrade cilium cilium/cilium --version 1.20.2 \
--namespace kube-system --reuse-values -f cilium-values.yaml
# The agents read the ConfigMap at start: restart them.
kubectl -n kube-system rollout restart daemonset/cilium
kubectl -n kube-system rollout status daemonset/cilium
# Connections opened before the change keep their pod-IP destination.
# Restart meshed workloads so every connection is re-opened through the sidecar.
kubectl -n app rollout restart deployment2. Kuma: strict mTLS, and no plaintext passthrough.
apiVersion: kuma.io/v1alpha1
kind: Mesh
metadata:
name: default
spec:
mtls:
enabledBackend: ca-1
backends:
- name: ca-1
type: builtin
mode: STRICT # servers reject plaintext (STRICT is also the default)
networking:
outbound:
# Unknown destinations are dropped, not forwarded in plaintext.
# Anything outside the mesh now needs a MeshPassthrough or a
# MeshExternalService entry. Check passthrough stats before you flip it.
passthrough: false
---
apiVersion: kuma.io/v1alpha1
kind: MeshTLS
metadata:
name: strict-everywhere
namespace: kuma-system
labels:
kuma.io/mesh: default
spec:
targetRef:
kind: Mesh
rules:
- default:
mode: Strict # every dataplane no narrower MeshTLS targets
tlsVersion:
min: TLS13
max: TLS13With mTLS on, Kuma denies traffic that no MeshTrafficPermission allows. Put your permissions in place first (see Default-deny MeshTrafficPermission).
3. Istio: strict mTLS mesh-wide.
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default
namespace: istio-system # the root namespace: applies to the whole mesh
spec:
mtls:
mode: STRICT # sidecars reject plaintext inbound4. Linkerd: refuse unauthenticated clients. mTLS survives socket load
balancing in Linkerd, but set hostNamespaceOnly anyway to keep its load
balancing, and stop accepting plaintext from unmeshed clients:
# linkerd-control-plane Helm values
proxy:
defaultInboundPolicy: all-authenticated # only meshed (mTLS) clientsProve it
Run on a throwaway kind cluster: Kubernetes 1.34.0, one control plane and two
workers, no kube-proxy, Cilium 1.19.3 with kubeProxyReplacement=true and
VXLAN tunneling, Kuma 2.14.3 with builtin mTLS. demo/web (nginx) runs on
lab-worker, demo/client (curl) on lab-worker2. The request carries a
fake password in the query string so it is easy to find on the wire.
1. The trap is on even when the ConfigMap says socket LB is off.
cilium config view | grep -E "^(kube-proxy-replacement|bpf-lb-sock) "
kubectl -n kube-system exec ds/cilium -c cilium-agent -- \
cilium-dbg status --verbose | grep -A3 "KubeProxyReplacement Details"bpf-lb-sock false
kube-proxy-replacement true
KubeProxyReplacement Details:
Status: True
Socket LB: Enabled
Socket LB Tracing: Enabled
Socket LB Coverage: Fullbpf-lb-sock is false in the ConfigMap, yet the agent runs socket load
balancing in every namespace (Coverage: Full), because kube-proxy
replacement turns it on. Trust the agent's status, not the ConfigMap.
2. Before the fix, mesh in STRICT mode. Send the request, and capture on the server's node:
kubectl -n demo exec deploy/client -c client -- curl -s -o /dev/null \
-w "HTTP %{http_code}\n" "http://web.demo.svc.cluster.local/?password=hunter2"
# on the server's node (in kind, the node is a container):
docker run --rm --net container:lab-worker nicolaka/netshoot:v0.14 \
tcpdump -i any -nn -A 'tcp port 80'HTTP 000
command terminated with exit code 52
20:39:44.251148 cilium_vxlan P IP 10.244.2.139.42496 > 10.244.1.152.80: Flags [P.], ... length 107: HTTP: GET /?password=hunter2 HTTP/1.1The server sidecar refused the plaintext, so curl got an empty reply (exit code 52). The password crossed the network in clear anyway.
3. Before the fix, mesh in PERMISSIVE mode. Same request, same capture:
HTTP 200
20:40:08.299917 cilium_vxlan P IP 10.244.2.139.43096 > 10.244.1.152.80: Flags [P.], ... length 107: HTTP: GET /?password=hunter2 HTTP/1.1Download the capture before the fix (pcap, 4 KB)
Everything works, and the password is on the wire. This is the green
dashboard. The capture above is a separate recording of this state (socket
load balancing in pods, PERMISSIVE, HTTP 200): open it in Wireshark and
search for hunter2.
4. Apply the fix with helm, restart the agents, restart the workloads.
helm upgrade cilium cilium/cilium -n kube-system --version 1.19.3 \
--reuse-values --set socketLB.hostNamespaceOnly=true
kubectl -n kube-system rollout restart ds/cilium
cilium config view | grep -E "^bpf-lb-sock-hostns-only "
kubectl -n kube-system exec ds/cilium -c cilium-agent -- \
cilium-dbg status --verbose | grep "Socket LB"
kubectl -n demo rollout restart deploy/web deploy/clientSTATUS: deployed
bpf-lb-sock-hostns-only true
Socket LB: Enabled
Socket LB Tracing: Enabled
Socket LB Coverage: Hostns-only5. After the fix. The same request returned HTTP 200 in PERMISSIVE and
in STRICT mode, and the capture held no plaintext password in either mode. The
first data packet to the web pod, printed with tcpdump -X:
20:44:18.184803 cilium_vxlan P IP 10.244.2.197.41568 > 10.244.1.250.80: Flags [P.], ... length 1581: HTTP
0x0030: d1ae ec12 1603 0106 2801 0006 2403 03aa ........(...$...
0x0090: 05cf 0000 0022 0020 0000 1d77 6562 5f64 .....".....web_d
0x00a0: 656d 6f5f 7376 635f 3830 7b6d 6573 683d emo_svc_80{mesh=
0x00b0: 6465 6661 756c 747d 0017 0000 ff01 0001 default}........
0x00c0: 0000 0a00 0600 0400 1d00 1700 0b00 0201 ................Download the capture after the fix (pcap, 15 KB)
This capture is a separate recording of the fixed state
(socketLB.hostNamespaceOnly=true, STRICT, HTTP 200); hunter2 does not
appear in it. 16 03 01 starts a TLS handshake record. The server name is Kuma's mTLS name
for the service, web_demo_svc_80{mesh=default}. tcpdump still labels port
80 as HTTP; the bytes are TLS. The supported groups offered are 0x001d
(x25519) and 0x0017 (secp256r1) only: a default Kuma mesh is not
post-quantum.
6. Count instead of read. The first filter matches only TLS handshake
records (content type 0x16), the second only a plaintext GET at the start
of the TCP payload. Live, on the server pod's interface or the server node:
tcpdump -i any -nn -c 5 'tcp port 80 and (tcp[((tcp[12:1] & 0xf0) >> 2):1] = 0x16)'
tcpdump -i any -nn -c 5 'tcp port 80 and tcp[((tcp[12:1] & 0xf0) >> 2):4] = 0x47455420'The same filters, run with tcpdump 4.99.5 on the two published captures:
for f in before-fix.pcap after-fix.pcap; do
echo "$f: TLS handshake packets $(tcpdump -nn -r $f 'tcp[((tcp[12:1] & 0xf0) >> 2):1] = 0x16' 2>/dev/null | wc -l), plaintext GET packets $(tcpdump -nn -r $f 'tcp[((tcp[12:1] & 0xf0) >> 2):4] = 0x47455420' 2>/dev/null | wc -l)"
donebefore-fix.pcap: TLS handshake packets 0, plaintext GET packets 2
after-fix.pcap: TLS handshake packets 8, plaintext GET packets 0Checks not run in the lab
These help on a real cluster. They were not run for this page, so there is no output to show.
hubble observe --namespace demo --type trace-sock --last 20after a request. What you should see: before the fix, apre-xlate-fwd TRACEDline from the client to the Service and apost-xlate-fwd TRANSLATEDline to the server pod (the socket rewrite). After the fix, notrace-sockevents from pods. Hubble cannot show encryption here: the port is the application port either way.kumactl inspect dataplane <client-pod>.demo --type=stats | grep -i passthrough. What you should see: before the fix, the passthrough cluster'supstream_cx_totalgrows with each new connection toweb; after the fix it stays flat. The Kuma docs ask you to check these counters before settingpassthrough: false.- For Istio,
kubectl -n demo exec deploy/client -c istio-proxy -- pilot-agent request GET stats | grep 'PassthroughCluster.upstream_cx_total'. What you should see: no growth when the client calls another meshed Service.
Mistakes people make
Trusting the green dashboard
A mesh dashboard shows that mTLS is configured and that certificates were
issued. It does not show what crossed the wire. In the lab, the PERMISSIVE
mesh returned HTTP 200 with the password in clear on the network.
Believing strict mode stops the leak
STRICT turns the bug into errors, which is better. It does not stop the first request from crossing the network in plaintext: the server sidecar rejects it only after it arrives. Fix the socket load balancing; strict mode is the alarm, not the lock.
Reading bpf-lb-sock false and relaxing
With kube-proxy replacement on, the ConfigMap can say bpf-lb-sock false
while the agent runs socket load balancing with Coverage: Full. Check
cilium-dbg status --verbose on each node.
Using cilium upgrade --reuse-values for the fix
In the lab, cilium upgrade --reuse-values --set socketLB.hostNamespaceOnly=true
failed with nil pointer evaluating interface {}.enabled and changed nothing:
coverage stayed Full. helm upgrade --reuse-values with the same value
worked. Whatever tool you use, confirm Coverage: Hostns-only afterwards.
Changing the setting and stopping there
bpf-lb-sock-hostns-only is read when the agent starts. Restart the Cilium
DaemonSet, confirm the coverage on every node, then restart the meshed
workloads. Connections opened before the change keep their pod-IP destination
until they close.
Capturing on the wrong interface
Inside the server pod, the sidecar talks to the application over a local
connection, and that leg is plaintext by design. A capture on lo inside the
pod shows readable HTTP even when the network leg is encrypted. Capture on the
pod's eth0 or on the node, as in the lab.
Forgetting the Gateway API side
In Cilium 1.15, a user reported that Cilium's own Gateway API gateways never
got an address with hostNamespaceOnly: true
(cilium/cilium#31274). The
same report notes that Kuma, Linkerd and Istio all require the setting. The
Cilium 1.20.2 chart now forces bpf-lb-sock-hostns-only: "true" whenever
gatewayAPI.enabled is set. On older versions, test your gateways after the
change.
Flipping it cluster-wide at once
The setting changes which BPF programs load. On Cilium 1.18.3 with one kernel, the socket programs failed the verifier and ClusterIP access from the host namespace stopped working (cilium/cilium#42659). On Cilium 1.10 and 1.11, users reported pods losing the API server after enabling it (cilium/cilium#17738). Roll it out on one node first, read the agent log, and test Service access from a pod and from the host.
Checklist
- Cilium Helm values set
socketLB.hostNamespaceOnly: truewheneverkubeProxyReplacement: trueorsocketLB.enabled: true. - Cilium Helm values set
cni.exclusive: falsewhen a mesh CNI is installed. - Every Cilium agent reports
Socket LB Coverage: Hostns-onlyincilium-dbg status --verbose, whatever the ConfigMap says aboutbpf-lb-sock. - Meshed workloads were restarted after the agents.
hubble observe --type trace-sockshows no translations from meshed pods.- Kuma Mesh runs
mode: STRICTand no MeshTLS setsPermissive. - Kuma
networking.outbound.passthroughisfalse, with MeshPassthrough or MeshExternalService entries for real external destinations. - Istio has a mesh-wide PeerAuthentication with
mode: STRICT. - Linkerd's default inbound policy is
all-authenticated. - A capture on the server pod's
eth0or the server node shows TLS records (16 03 01) and no readable request line. - Passthrough counters stay flat for traffic between meshed services.
Cilium and your mesh each did their job correctly; they just did it in the wrong order. One Helm value puts them back in line, and one tcpdump tells you it worked.
H2-CSPE
Learn it on a live range
Service mesh and gateways, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Dome
Want it run for you?
The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.
See the Dome