Service mesh

Cilium with a sidecar mesh (Kuma, Istio, Linkerd) without losing mTLS, the secure way

We know why you're here. Kuma swears your traffic is mTLS, the dashboard is green, and tcpdump just showed you a password in plain text. You're not crazy, you heard right: Cilium did it, very politely, before your sidecar ever saw the packet.

The short answer

When Cilium replaces kube-proxy, set socketLB.hostNamespaceOnly=true so connections inside pods keep the Service ClusterIP and the sidecar can match it. Set cni.exclusive=false for the mesh CNI. Run the mesh in strict mTLS and turn off plaintext passthrough so a bypass fails closed. Then prove encryption on the wire with tcpdump and proxy stats.

Updated Houssam Hammoudi, CTOTested with Kubernetes 1.34.0 (kind), Cilium 1.19.3, Kuma 2.14.3; configuration also checked against the Cilium 1.20.2 chart and Kuma 2.14.5 CRDs

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Cilium's kube-proxy replacement load-balances Services at the socket. A BPF program attached to the cgroup runs on connect() (and sendmsg() for UDP). When an application connects to a ClusterIP, the program picks a backend and rewrites the destination to that backend's pod IP. This happens before the first packet exists.

A sidecar mesh works one step later. Inside the pod, iptables rules (from an init container or the mesh CNI) redirect outbound TCP to the sidecar. The sidecar reads the original destination and matches it against the Services it knows, by ClusterIP and port. That match selects the upstream, and the upstream carries the mTLS settings.

After the socket rewrite, the original destination is a pod IP, not the ClusterIP. The sidecar has no listener for that address. Kuma documents what happens next: an unknown destination goes through passthrough, and passthrough is not encrypted. The connection leaves the pod as plain TCP. Telemetry, retries, timeouts and any policy that needs the client's certificate never see it either.

What happens at the server depends on its mTLS mode:

  • Strict: the server sidecar expects TLS, receives plaintext, and closes the connection. The client gets an empty reply. But the request, with any password or token in it, has already crossed the network in clear. The usual "fix" for the errors is to switch to permissive.
  • Permissive: the server sidecar accepts plaintext. Everything works, the mesh dashboard stays green, and anyone who can capture on a node, a hypervisor or the network between nodes can read the traffic.

Linkerd is the exception worth knowing. Its proxy also resolves pod IPs, so Linkerd's docs say mTLS and telemetry keep working. What you lose there is Linkerd's load balancing and dynamic request routing.

Nothing logs an error. The trap is on whenever socket load balancing runs in pod namespaces: with kubeProxyReplacement: true, or with socketLB.enabled: true next to kube-proxy.

What the docs say

These settings prevent Cilium’s socket-based load balancing from interfering with Istio’s proxying.

Source: Cilium docs, Integration with Istio

Without this option, when Cilium does service resolution via socket load balancing, Istio sidecar will be bypassed, resulting in loss of Istio features including encryption and telemetry.

Source: Cilium 1.13 docs, Getting Started Using Istio

If you pass requests to direct IP addresses, Envoy considers them unknown destinations and manages them in passthrough mode – which means they’re not encrypted with mTLS.

Source: Kuma docs, Kubernetes annotations and labels (kuma.io/direct-access-services)

Consequentially, while mTLS and telemetry will still function correctly, features such as peak EWMA load balancing, and dynamic request routing may not work as expected.

Source: Linkerd docs, Cluster Configuration (Cilium)

Cilium names the setting for Istio only, and the sentence that says "encryption" was dropped from the current page; it survives in the 1.13 docs. Kuma's transparent proxying page does not mention Cilium at all, and Kuma's CNI page only tells Cilium users to set cni.exclusive to false. Kuma explains the plaintext passthrough under an unrelated annotation. No page shows how to check the wire.

The secure configuration

1. Cilium: keep socket load balancing in the host namespace.

yaml
# cilium-values.yaml (Helm)
kubeProxyReplacement: true
k8sServiceHost: 192.0.2.10        # your API server address
k8sServicePort: 6443
socketLB:
  # Socket LB stays in the host namespace. Inside pods, connect() keeps the
  # ClusterIP, so the sidecar can match the Service. Pod traffic that is not
  # meshed is still load-balanced by Cilium's tc program on the veth.
  # Renders as bpf-lb-sock-hostns-only: "true" in the cilium-config ConfigMap.
  hostNamespaceOnly: true
cni:
  # Do not delete other CNI configs: kuma-cni, istio-cni and linkerd-cni
  # chain after Cilium. Renders as cni-exclusive: "false".
  exclusive: false
bash
# Use helm, not "cilium upgrade --reuse-values" (see Mistakes).
helm upgrade cilium cilium/cilium --version 1.20.2 \
  --namespace kube-system --reuse-values -f cilium-values.yaml

# The agents read the ConfigMap at start: restart them.
kubectl -n kube-system rollout restart daemonset/cilium
kubectl -n kube-system rollout status daemonset/cilium

# Connections opened before the change keep their pod-IP destination.
# Restart meshed workloads so every connection is re-opened through the sidecar.
kubectl -n app rollout restart deployment

2. Kuma: strict mTLS, and no plaintext passthrough.

yaml
apiVersion: kuma.io/v1alpha1
kind: Mesh
metadata:
  name: default
spec:
  mtls:
    enabledBackend: ca-1
    backends:
      - name: ca-1
        type: builtin
        mode: STRICT            # servers reject plaintext (STRICT is also the default)
  networking:
    outbound:
      # Unknown destinations are dropped, not forwarded in plaintext.
      # Anything outside the mesh now needs a MeshPassthrough or a
      # MeshExternalService entry. Check passthrough stats before you flip it.
      passthrough: false
---
apiVersion: kuma.io/v1alpha1
kind: MeshTLS
metadata:
  name: strict-everywhere
  namespace: kuma-system
  labels:
    kuma.io/mesh: default
spec:
  targetRef:
    kind: Mesh
  rules:
    - default:
        mode: Strict            # every dataplane no narrower MeshTLS targets
        tlsVersion:
          min: TLS13
          max: TLS13

With mTLS on, Kuma denies traffic that no MeshTrafficPermission allows. Put your permissions in place first (see Default-deny MeshTrafficPermission).

3. Istio: strict mTLS mesh-wide.

yaml
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
  name: default
  namespace: istio-system       # the root namespace: applies to the whole mesh
spec:
  mtls:
    mode: STRICT                # sidecars reject plaintext inbound

4. Linkerd: refuse unauthenticated clients. mTLS survives socket load balancing in Linkerd, but set hostNamespaceOnly anyway to keep its load balancing, and stop accepting plaintext from unmeshed clients:

yaml
# linkerd-control-plane Helm values
proxy:
  defaultInboundPolicy: all-authenticated   # only meshed (mTLS) clients

Prove it

Run on a throwaway kind cluster: Kubernetes 1.34.0, one control plane and two workers, no kube-proxy, Cilium 1.19.3 with kubeProxyReplacement=true and VXLAN tunneling, Kuma 2.14.3 with builtin mTLS. demo/web (nginx) runs on lab-worker, demo/client (curl) on lab-worker2. The request carries a fake password in the query string so it is easy to find on the wire.

1. The trap is on even when the ConfigMap says socket LB is off.

bash
cilium config view | grep -E "^(kube-proxy-replacement|bpf-lb-sock) "
kubectl -n kube-system exec ds/cilium -c cilium-agent -- \
  cilium-dbg status --verbose | grep -A3 "KubeProxyReplacement Details"
text
bpf-lb-sock                                       false
kube-proxy-replacement                            true
KubeProxyReplacement Details:
  Status:               True
  Socket LB:            Enabled
  Socket LB Tracing:    Enabled
  Socket LB Coverage:   Full

bpf-lb-sock is false in the ConfigMap, yet the agent runs socket load balancing in every namespace (Coverage: Full), because kube-proxy replacement turns it on. Trust the agent's status, not the ConfigMap.

2. Before the fix, mesh in STRICT mode. Send the request, and capture on the server's node:

bash
kubectl -n demo exec deploy/client -c client -- curl -s -o /dev/null \
  -w "HTTP %{http_code}\n" "http://web.demo.svc.cluster.local/?password=hunter2"
# on the server's node (in kind, the node is a container):
docker run --rm --net container:lab-worker nicolaka/netshoot:v0.14 \
  tcpdump -i any -nn -A 'tcp port 80'
text
HTTP 000
command terminated with exit code 52
20:39:44.251148 cilium_vxlan P   IP 10.244.2.139.42496 > 10.244.1.152.80: Flags [P.], ... length 107: HTTP: GET /?password=hunter2 HTTP/1.1

The server sidecar refused the plaintext, so curl got an empty reply (exit code 52). The password crossed the network in clear anyway.

3. Before the fix, mesh in PERMISSIVE mode. Same request, same capture:

text
HTTP 200
20:40:08.299917 cilium_vxlan P   IP 10.244.2.139.43096 > 10.244.1.152.80: Flags [P.], ... length 107: HTTP: GET /?password=hunter2 HTTP/1.1

Download the capture before the fix (pcap, 4 KB)

Everything works, and the password is on the wire. This is the green dashboard. The capture above is a separate recording of this state (socket load balancing in pods, PERMISSIVE, HTTP 200): open it in Wireshark and search for hunter2.

4. Apply the fix with helm, restart the agents, restart the workloads.

bash
helm upgrade cilium cilium/cilium -n kube-system --version 1.19.3 \
  --reuse-values --set socketLB.hostNamespaceOnly=true
kubectl -n kube-system rollout restart ds/cilium
cilium config view | grep -E "^bpf-lb-sock-hostns-only "
kubectl -n kube-system exec ds/cilium -c cilium-agent -- \
  cilium-dbg status --verbose | grep "Socket LB"
kubectl -n demo rollout restart deploy/web deploy/client
text
STATUS: deployed
bpf-lb-sock-hostns-only                           true
  Socket LB:            Enabled
  Socket LB Tracing:    Enabled
  Socket LB Coverage:   Hostns-only

5. After the fix. The same request returned HTTP 200 in PERMISSIVE and in STRICT mode, and the capture held no plaintext password in either mode. The first data packet to the web pod, printed with tcpdump -X:

text
20:44:18.184803 cilium_vxlan P   IP 10.244.2.197.41568 > 10.244.1.250.80: Flags [P.], ... length 1581: HTTP
	0x0030:  d1ae ec12 1603 0106 2801 0006 2403 03aa  ........(...$...
	0x0090:  05cf 0000 0022 0020 0000 1d77 6562 5f64  .....".....web_d
	0x00a0:  656d 6f5f 7376 635f 3830 7b6d 6573 683d  emo_svc_80{mesh=
	0x00b0:  6465 6661 756c 747d 0017 0000 ff01 0001  default}........
	0x00c0:  0000 0a00 0600 0400 1d00 1700 0b00 0201  ................

Download the capture after the fix (pcap, 15 KB)

This capture is a separate recording of the fixed state (socketLB.hostNamespaceOnly=true, STRICT, HTTP 200); hunter2 does not appear in it. 16 03 01 starts a TLS handshake record. The server name is Kuma's mTLS name for the service, web_demo_svc_80{mesh=default}. tcpdump still labels port 80 as HTTP; the bytes are TLS. The supported groups offered are 0x001d (x25519) and 0x0017 (secp256r1) only: a default Kuma mesh is not post-quantum.

6. Count instead of read. The first filter matches only TLS handshake records (content type 0x16), the second only a plaintext GET at the start of the TCP payload. Live, on the server pod's interface or the server node:

bash
tcpdump -i any -nn -c 5 'tcp port 80 and (tcp[((tcp[12:1] & 0xf0) >> 2):1] = 0x16)'
tcpdump -i any -nn -c 5 'tcp port 80 and tcp[((tcp[12:1] & 0xf0) >> 2):4] = 0x47455420'

The same filters, run with tcpdump 4.99.5 on the two published captures:

bash
for f in before-fix.pcap after-fix.pcap; do
  echo "$f: TLS handshake packets $(tcpdump -nn -r $f 'tcp[((tcp[12:1] & 0xf0) >> 2):1] = 0x16' 2>/dev/null | wc -l), plaintext GET packets $(tcpdump -nn -r $f 'tcp[((tcp[12:1] & 0xf0) >> 2):4] = 0x47455420' 2>/dev/null | wc -l)"
done
text
before-fix.pcap: TLS handshake packets 0, plaintext GET packets 2
after-fix.pcap: TLS handshake packets 8, plaintext GET packets 0

Checks not run in the lab

These help on a real cluster. They were not run for this page, so there is no output to show.

  • hubble observe --namespace demo --type trace-sock --last 20 after a request. What you should see: before the fix, a pre-xlate-fwd TRACED line from the client to the Service and a post-xlate-fwd TRANSLATED line to the server pod (the socket rewrite). After the fix, no trace-sock events from pods. Hubble cannot show encryption here: the port is the application port either way.
  • kumactl inspect dataplane <client-pod>.demo --type=stats | grep -i passthrough. What you should see: before the fix, the passthrough cluster's upstream_cx_total grows with each new connection to web; after the fix it stays flat. The Kuma docs ask you to check these counters before setting passthrough: false.
  • For Istio, kubectl -n demo exec deploy/client -c istio-proxy -- pilot-agent request GET stats | grep 'PassthroughCluster.upstream_cx_total'. What you should see: no growth when the client calls another meshed Service.

Mistakes people make

Trusting the green dashboard

A mesh dashboard shows that mTLS is configured and that certificates were issued. It does not show what crossed the wire. In the lab, the PERMISSIVE mesh returned HTTP 200 with the password in clear on the network.

Believing strict mode stops the leak

STRICT turns the bug into errors, which is better. It does not stop the first request from crossing the network in plaintext: the server sidecar rejects it only after it arrives. Fix the socket load balancing; strict mode is the alarm, not the lock.

Reading bpf-lb-sock false and relaxing

With kube-proxy replacement on, the ConfigMap can say bpf-lb-sock false while the agent runs socket load balancing with Coverage: Full. Check cilium-dbg status --verbose on each node.

Using cilium upgrade --reuse-values for the fix

In the lab, cilium upgrade --reuse-values --set socketLB.hostNamespaceOnly=true failed with nil pointer evaluating interface {}.enabled and changed nothing: coverage stayed Full. helm upgrade --reuse-values with the same value worked. Whatever tool you use, confirm Coverage: Hostns-only afterwards.

Changing the setting and stopping there

bpf-lb-sock-hostns-only is read when the agent starts. Restart the Cilium DaemonSet, confirm the coverage on every node, then restart the meshed workloads. Connections opened before the change keep their pod-IP destination until they close.

Capturing on the wrong interface

Inside the server pod, the sidecar talks to the application over a local connection, and that leg is plaintext by design. A capture on lo inside the pod shows readable HTTP even when the network leg is encrypted. Capture on the pod's eth0 or on the node, as in the lab.

Forgetting the Gateway API side

In Cilium 1.15, a user reported that Cilium's own Gateway API gateways never got an address with hostNamespaceOnly: true (cilium/cilium#31274). The same report notes that Kuma, Linkerd and Istio all require the setting. The Cilium 1.20.2 chart now forces bpf-lb-sock-hostns-only: "true" whenever gatewayAPI.enabled is set. On older versions, test your gateways after the change.

Flipping it cluster-wide at once

The setting changes which BPF programs load. On Cilium 1.18.3 with one kernel, the socket programs failed the verifier and ClusterIP access from the host namespace stopped working (cilium/cilium#42659). On Cilium 1.10 and 1.11, users reported pods losing the API server after enabling it (cilium/cilium#17738). Roll it out on one node first, read the agent log, and test Service access from a pod and from the host.

Checklist

  • Cilium Helm values set socketLB.hostNamespaceOnly: true whenever kubeProxyReplacement: true or socketLB.enabled: true.
  • Cilium Helm values set cni.exclusive: false when a mesh CNI is installed.
  • Every Cilium agent reports Socket LB Coverage: Hostns-only in cilium-dbg status --verbose, whatever the ConfigMap says about bpf-lb-sock.
  • Meshed workloads were restarted after the agents.
  • hubble observe --type trace-sock shows no translations from meshed pods.
  • Kuma Mesh runs mode: STRICT and no MeshTLS sets Permissive.
  • Kuma networking.outbound.passthrough is false, with MeshPassthrough or MeshExternalService entries for real external destinations.
  • Istio has a mesh-wide PeerAuthentication with mode: STRICT.
  • Linkerd's default inbound policy is all-authenticated.
  • A capture on the server pod's eth0 or the server node shows TLS records (16 03 01) and no readable request line.
  • Passthrough counters stay flat for traffic between meshed services.

Cilium and your mesh each did their job correctly; they just did it in the wrong order. One Helm value puts them back in line, and one tcpdump tells you it worked.

H2-CSPE

Learn it on a live range

Service mesh and gateways, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Dome

Want it run for you?

The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.

See the Dome