Service mesh

Mesh sidecars inside gVisor, the secure way

You put the untrusted workload in gVisor and in the mesh, and got two outcomes: the init container crashes on iptables, or everything starts and you are not sure the sidecar sees a single packet. Both are the same question: who programs the network, and in which kernel?

The short answer

In a gVisor pod the sidecar and its traffic rules live inside the sandbox's own network stack. Use the mesh's init container, not its CNI plugin, and a dedicated runsc handler with net-raw enabled so that init container can set iptables. Keep net-raw off everywhere else, drop all capabilities from the app, and verify TLS on the wire.

Updated Houssam Hammoudi, CTOTested with Kuma 2.14.3 and gVisor release-20260921 on Kubernetes 1.34.0 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A sidecar mesh intercepts traffic with iptables rules in the pod's network namespace. The rules come from an init container inside the pod or from the mesh's CNI plugin on the node. Either way, they assume the pod's packets go through the Linux kernel's network stack.

gVisor does not work that way. It runs its own network stack, netstack, inside the sandbox. The application, the sidecar and any iptables rules for them all live in that user-space stack. gVisor writes data link layer packets directly to the pod's virtual network device.

That breaks meshes in two directions:

  • Init container mode fails loudly. The init container calls iptables inside the sandbox. A gVisor maintainer states that Istio works when runsc runs with --net-raw=true, and lists the supported parts of iptables as the filter and nat tables and the PREROUTING, INPUT and OUTPUT chains. Without --net-raw, the init container cannot set its rules and the pod does not start. Kuma is not named; if its init container fails, its log names the rule gVisor rejected.
  • CNI mode can fail quietly. A mesh CNI plugin writes its rules into the pod's network namespace in the host kernel. A gVisor maintainer explains that Istio CNI does not work for this reason: the network is virtualized inside the sandbox, so rules changed on the host "will not have the desired effect". The pod can start, the sidecar can look healthy, and traffic can leave without it. Without the sidecar there is no mTLS, as in Cilium with a sidecar mesh.

The tempting fixes are the dangerous ones: --network=host, which hands the sandbox the host's network stack, or turning on --net-raw for every gVisor pod on the node.

What the docs say

All aspects of the network stack are handled inside the Sentry — including TCP connection state, control messages, and packet assembly — keeping it isolated from the host network stack.

Source: gVisor docs, Networking

iptables are only partially supported.

Source: gVisor docs, Compatibility

Heads up: Istio is supported in gVisor. Make sure to run with runsc flag --net-raw=true (required to setup the iptables rules Istio relies on).

Source: gVisor maintainer, google/gvisor#170

Raw sockets allow malicious containers to craft packets and potentially attack the network.

Source: gVisor source, runsc/config/flags.go (net-raw flag help)

gVisor's user guide does not have a service mesh page, and a search of the Kuma and Istio docs finds no mention of gVisor. The answer is spread across issue comments and a flag's help text, and the help text is the warning: the flag that makes the mesh work is the one gVisor keeps off by default for safety.

The secure configuration

1. Two runsc handlers on sandbox nodes. Keep the default handler without raw sockets, and add a second one only for meshed sandboxed pods.

toml
# /etc/containerd/config.toml (fragment)
version = 2
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runsc]
  runtime_type = "io.containerd.runsc.v1"
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runsc.options]
  TypeUrl = "io.containerd.runsc.v1.options"
  ConfigPath = "/etc/containerd/runsc.toml"
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runsc-mesh]
  runtime_type = "io.containerd.runsc.v1"
[plugins."io.containerd.grpc.v1.cri".containerd.runtimes.runsc-mesh.options]
  TypeUrl = "io.containerd.runsc.v1.options"
  ConfigPath = "/etc/containerd/runsc-mesh.toml"
toml
# /etc/containerd/runsc.toml: every other gVisor pod keeps raw sockets off.
[runsc_config]
  network = "sandbox"
toml
# /etc/containerd/runsc-mesh.toml: only this handler allows raw sockets,
# which the mesh init container needs to program the sandbox's iptables.
[runsc_config]
  network = "sandbox"   # gVisor netstack, never "host"
  net-raw = "true"

This is the containerd 1.x (version = 2) layout from the gVisor containerd docs; on containerd 2.x, the gVisor docs point to the version = 3 header, which changes the plugin path. Values under [runsc_config] become runsc flags (net-raw = "true" becomes --net-raw="true"). Restart containerd on the node after the change. Releases since 2026-07 ship runsc, containerd-shim-runsc-v1 and a gvisor-bin/ directory of helper binaries that runsc looks for next to itself; install all of them.

2. A RuntimeClass for meshed sandboxes.

yaml
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
  name: gvisor-mesh
handler: runsc-mesh
scheduling:
  nodeSelector:
    sandbox.example.com/gvisor: "true"
  tolerations:                       # the gVisor pool is tainted; without this, pods never schedule
    - key: sandbox.example.com/gvisor
      operator: Exists
      effect: NoSchedule

3. Use the mesh's init container for these pods. With Kuma, that is the default when the Kuma CNI is not installed (cni.enabled: false, the Helm chart default); with Istio, the istio-init container is the default without the Istio CNI node agent. If your cluster runs the mesh CNI, keep sandboxed pods off it or prove interception as shown below before trusting it.

4. The workload: no capabilities, gVisor, mesh.

yaml
# Kuma ContainerPatch: what gVisor's netstack needs from the mesh init container,
# and the one place the dataplane token is mounted.
apiVersion: kuma.io/v1alpha1
kind: ContainerPatch
metadata:
  name: gvisor-init
  namespace: kuma-system
spec:
  initPatch:
    - op: add
      path: /args/-
      value: '"--disable-comments"'               # gVisor's iptables has no comment match
    - op: add
      path: /args/-
      value: '"--skip-dns-conntrack-zone-split"'  # gVisor's iptables has no CT target
  sidecarPatch:
    - op: add
      path: /volumeMounts/-
      value: '{"name": "kuma-token", "mountPath": "/var/run/secrets/kubernetes.io/serviceaccount", "readOnly": true}'
yaml
apiVersion: v1
kind: Pod
metadata:
  name: untrusted-worker
  namespace: sandbox          # labelled for sidecar injection; see Pod Security below
  labels:
    app: untrusted-worker
  annotations:
    kuma.io/container-patches: gvisor-init
    kuma.io/builtin-dns: disabled     # gVisor does not redirect DNS to the sidecar's resolver
spec:
  runtimeClassName: gvisor-mesh
  automountServiceAccountToken: false # the untrusted app gets no token
  containers:
    - name: app
      image: registry.example.com/untrusted-worker@sha256:0000000000000000000000000000000000000000000000000000000000000000
      securityContext:
        allowPrivilegeEscalation: false
        runAsNonRoot: true
        runAsUser: 10001
        capabilities:
          drop: ["ALL"]            # no NET_RAW for the workload, even though the sandbox allows it
        seccompProfile:
          type: RuntimeDefault
  volumes:
    - name: kuma-token             # the sidecar's identity; mounted into the sidecar only, by the patch
      projected:
        sources:
          - serviceAccountToken:
              path: token
              expirationSeconds: 3600

The injected init container (kuma-init or istio-init) asks for NET_ADMIN and NET_RAW. Inside gVisor those act on the sandbox's own network stack, not the node's. Pod Security baseline still rejects them, so the namespace needs an exemption scoped to the mesh init container (for example with a ValidatingAdmissionPolicy), not privileged for everything.

Prove it

Run on a lab cluster (Kubernetes 1.34, Kuma 2.14.3, gVisor release-20260921) with the containerd fragment, RuntimeClass, ContainerPatch and Pod above; the app image placeholder was replaced by busybox:1.37, and the namespace was exempt from Pod Security baseline (see the mistakes). Real output.

1. The pod runs on gVisor, meshed:

bash
kubectl -n sandbox get pod untrusted-worker
kubectl -n sandbox exec untrusted-worker -c app -- dmesg | head -2
text
untrusted-worker   2/2   Running   0     12s
[   0.000000] Starting gVisor...
[   0.591325] Recruiting cron-ies...

2. The mesh init container finished and the proxy is online:

bash
kubectl -n sandbox get pod untrusted-worker \
  -o jsonpath='{range .status.initContainerStatuses[*]}{.name}{" "}{.state.terminated.reason}{"\n"}{end}'
kumactl inspect dataplanes --mesh default | grep untrusted-worker
text
kuma-init Completed
kuma-sidecar
untrusted-worker.sandbox ... Online ...

3. The untrusted app holds no token; only the sidecar does:

bash
kubectl -n sandbox exec untrusted-worker -c app -- ls /var/run/secrets/kubernetes.io/serviceaccount
text
ls: /var/run/secrets/kubernetes.io/serviceaccount: No such file or directory

4. Its traffic leaves as mTLS. The app calls another meshed service with ?password=hunter2 in the URL while tcpdump runs on the node:

text
backend
packets: 134, TLS handshakes from pod: 4, hunter2 visible on the wire: 0

Download the capture (pcap, 25 KB)

5. Only the pods you expect use the raw-socket handler:

bash
kubectl get runtimeclass -o custom-columns=NAME:.metadata.name,HANDLER:.handler
text
NAME           HANDLER
gvisor         runsc
gvisor-mesh    runsc-mesh
gvisor-nonet   runsc-nonet

Mistakes people make

Using the default mesh init container

Kuma's default rules use two iptables features gVisor's netstack does not have. The init container retries five times, then the pod loops in Init:CrashLoopBackOff:

text
restoring failed: exit status 1: Warning: Extension comment revision 0 not supported, missing kernel module?, Warning: Extension CT is not supported, missing kernel module?, ... iptables-restore: line 9 failed
Error: failed to setup transparent proxy: unable to restore iptables rules: /usr/sbin/iptables-legacy-restore failed

The two kuma-init flags in the ContainerPatch, --disable-comments and --skip-dns-conntrack-zone-split, remove both. The per-pod annotation traffic.kuma.io/transparent-proxy-config is not read by the 2.14 injector; on the lab it changed nothing.

Keeping the mesh's DNS

With DNS redirected to the sidecar, every name lookup inside gVisor failed (wget: bad address 'backend.app.svc.cluster.local:8080'). With kuma.io/builtin-dns: disabled, the app uses cluster DNS and the traffic is still intercepted and encrypted.

Taking the token away from the whole pod

automountServiceAccountToken: false is right for the app, but the sidecar proves its identity with the ServiceAccount token and crash-loops without it:

text
Error: dataplane token is invalid, in Kubernetes you must mount a serviceAccount token ...

Shadowing the token path with an empty volume in the app container does not work either: the injector copies the app's mount into the sidecar. Mount a projected token into the sidecar only, as the ContainerPatch does.

A RuntimeClass that ignores the pool's taint

If the gVisor pool is tainted, the RuntimeClass needs the toleration; without it the pod stays Pending with untolerated taint {sandbox.example.com/gvisor: true}. The same goes for any admission policy that lists allowed gVisor classes: add gvisor-mesh to it.

Setting --network=host to make it work

It removes gVisor's network isolation, which for many workloads is the point of running gVisor. The gVisor FAQ describes it as less secure. Fix the interception instead.

Enabling net-raw for every sandbox

--net-raw on the default handler gives raw sockets to every gVisor pod on the node. Put it on a separate handler and RuntimeClass used only by meshed sandboxed pods.

Trusting a healthy sidecar

A running sidecar proves nothing about interception. In CNI mode the pod can be green while its traffic goes around the proxy. Check the wire.

Leaving NET_RAW on the workload

With net-raw on, any container in the sandbox that keeps CAP_NET_RAW can craft packets on the pod's network. Drop all capabilities from application containers.

Copying only the runsc binary

Releases since 2026-07 need the gvisor-bin/ helper directory next to runsc. The gVisor install docs say runsc install tries to download missing helpers as a stopgap, and that this fallback ends at the end of September 2026. Install from the release tarball or the apt repository.

Checklist

  • Sandbox nodes have two runsc handlers: default without net-raw, and runsc-mesh with net-raw = "true".
  • Both handlers use network = "sandbox".
  • A gvisor-mesh RuntimeClass maps to runsc-mesh and is used only by meshed sandboxed pods.
  • Meshed sandboxed pods use the mesh init container, or interception in CNI mode was proven on the wire.
  • Application containers drop all capabilities.
  • The Pod Security exemption covers the mesh init container only.
  • dmesg in the pod shows gVisor.
  • A capture on the node shows TLS, not readable HTTP, from sandboxed pods.

gVisor gives the workload its own kernel, network stack included. The mesh has to be set up inside that stack, and the only honest check is the one on the wire.

H2-CSPE

Learn it on a live range

Service mesh and gateways, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Dome

Want it run for you?

The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.

See the Dome