Service mesh

Envoy Gateway as a delegated mesh gateway, the secure way

Envoy Gateway terminates TLS at the edge, Kuma encrypts everything inside, and the hop between them, from the gateway to your first service, is plain HTTP to a pod IP that the mesh has never heard of.

The short answer

Inject a Kuma sidecar into the Envoy Gateway proxy pods only, with the kuma.io/gateway: enabled annotation, and set routingType: Service in the EnvoyProxy resource so the gateway connects to Service ClusterIPs the sidecar can match. Label the namespace disabled so the pod label is honoured, mount a token into the sidecar only, exclude the xDS port, keep the mesh in strict mTLS, and allow backends only from the gateway's service.

Updated Houssam Hammoudi, CTOTested with Envoy Gateway v1.9.1, Gateway API v1, Kuma 2.14.3 (builtin CA, strict mTLS, passthrough off), Cilium 1.20.2, Kubernetes 1.34 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A delegated gateway is an existing API gateway that Kuma adopts: Kuma adds a sidecar to the gateway pod, leaves incoming traffic to the gateway, and takes over when the gateway calls into the mesh. The mesh then encrypts that hop with mTLS and applies its permissions.

This depends on one detail. The sidecar recognizes mesh traffic by the destination address: a Service ClusterIP and port it knows. Envoy Gateway, by default, load-balances itself and connects straight to pod IPs. The sidecar sees an unknown address and handles it in passthrough mode, which is not mTLS. It is the same failure as Cilium socket load balancing with a sidecar mesh, caused by a different component.

The result is one of two things:

  • with the mesh in permissive mode, the gateway's requests reach services in plaintext, with no mesh identity, and mesh permissions cannot tell the gateway from anyone else;
  • with strict mode, the requests fail with a 503, and the next change is usually to loosen the mesh.

Two more things stop the gateway from joining the mesh at all. Kuma only injects a pod by its own label when the pod's namespace also carries the injection label (any value). And Envoy Gateway turns off the service account token on its proxy pods, which the Kuma sidecar needs to register.

Kuma's delegated gateway docs solve the routing problem for ingress controllers with the ingress.kubernetes.io/service-upstream annotation. Envoy Gateway's docs do not list that annotation; the Envoy Gateway setting with the same effect is the routingType field.

What the docs say

In delegated gateway mode, Kuma configures an Envoy sidecar for your API gateway.

Source: Kuma docs, Delegated gateways

With this annotation the ingress controller sends traffic to the Service IP instead of directly to the endpoints selected by the Service.

Source: Kuma docs, Delegated gateways

RoutingType can be set to “Service” to use the Service Cluster IP for routing to the backend, or it can be set to “Endpoint” to use Endpoint routing. The default is “Endpoint”.

Source: Envoy Gateway docs, API reference (EnvoyProxySpec)

If you pass requests to direct IP addresses, Envoy considers them unknown destinations and manages them in passthrough mode – which means they’re not encrypted with mTLS.

Source: Kuma docs, Kubernetes annotations and labels

Labeling pods or deployments will take precedence on the namespace annotation.

Source: Kuma docs, Kubernetes annotations and labels

Kuma documents the routing fix in ingress-controller terms, and its delegated gateway page names Kong, not Envoy Gateway. Envoy Gateway's API reference describes routingType as a plain routing option and does not mention a service mesh. Nobody tells you that the Envoy Gateway default is the one that breaks mTLS.

The label sentence is true only in part. Kuma's pod injector webhook has this namespace selector (from the lab cluster):

text
{"key":"kuma.io/sidecar-injection","operator":"Exists"}

A pod label in a namespace without that label key never reaches the webhook. The pod label wins over the namespace label only when both exist.

The secure configuration

1. Label the proxy namespace disabled. This keeps the Envoy Gateway controller out of the mesh and lets the pod label below take effect.

bash
kubectl label namespace envoy-gateway-system kuma.io/sidecar-injection=disabled

2. An EnvoyProxy that makes the proxy pods delegated gateways.

yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: EnvoyProxy
metadata:
  name: mesh-delegated
  namespace: envoy-gateway-system
spec:
  routingType: Service              # send to the Service ClusterIP, so the Kuma sidecar matches it
  provider:
    type: Kubernetes
    kubernetes:
      envoyDeployment:
        pod:
          labels:
            kuma.io/sidecar-injection: enabled   # inject only the proxy pods, not the controller
          annotations:
            kuma.io/gateway: enabled             # delegated gateway: no inbound listeners
            traffic.kuma.io/exclude-outbound-ports: "18000"   # xDS to the Envoy Gateway controller
            kuma.io/container-patches: gateway-token          # token for the Kuma sidecar only
          volumes:
            - name: kuma-token
              projected:
                sources:
                  - serviceAccountToken:
                      path: token
                      expirationSeconds: 3600
---
apiVersion: kuma.io/v1alpha1
kind: ContainerPatch
metadata:
  name: gateway-token
  namespace: kuma-system            # ContainerPatches live in the Kuma system namespace
spec:
  sidecarPatch:
    - op: add
      path: /volumeMounts/-
      value: '{"name": "kuma-token", "mountPath": "/var/run/secrets/kubernetes.io/serviceaccount", "readOnly": true}'

Envoy Gateway sets automountServiceAccountToken: false on its proxy pods. Leave it that way: the envoy container faces the internet and has no use for a Kubernetes API token. The projected volume and the ContainerPatch give the token to kuma-sidecar alone.

The proxies fetch their configuration (xDS) from the envoy-gateway Service on port 18000, the controller's default xDS port. The controller is not in the mesh, so with passthrough off that connection fails unless the port is excluded. The Helm chart also exposes 18001 (ratelimit, used by the rate limit service, not the proxies) and 18002 (wasm, from which proxies download Wasm extensions). If you use Wasm extensions, exclude 18002 as well: "18000,18002".

3. Use it from the GatewayClass, and keep route attachment narrow.

yaml
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata:
  name: eg-mesh
spec:
  controllerName: gateway.envoyproxy.io/gatewayclass-controller
  parametersRef:
    group: gateway.envoyproxy.io
    kind: EnvoyProxy
    name: mesh-delegated
    namespace: envoy-gateway-system
---
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: edge
  namespace: edge
spec:
  gatewayClassName: eg-mesh
  listeners:
    - name: https
      protocol: HTTPS
      port: 443
      hostname: shop.example.com
      tls:
        mode: Terminate
        certificateRefs:
          - kind: Secret
            name: shop-example-com
      allowedRoutes:
        namespaces:
          from: Selector
          selector:
            matchLabels:
              edge.example.com/routes: allowed
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: shop
  namespace: shop                   # labelled edge.example.com/routes=allowed
spec:
  parentRefs:
    - name: edge
      namespace: edge
  hostnames:
    - shop.example.com
  rules:
    - backendRefs:
        - name: storefront
          port: 8080

4. The mesh: strict mTLS, no passthrough, and a permission for the gateway. Find the gateway's service name as Kuma sees it:

bash
kubectl -n envoy-gateway-system get dataplanes \
  -o custom-columns=NAME:.metadata.name,SERVICE:'.spec.networking.gateway.tags.kuma\.io/service'

Then allow only that service to call the backend:

yaml
apiVersion: kuma.io/v1alpha1
kind: MeshTrafficPermission
metadata:
  name: storefront-from-gateway
  namespace: shop
  labels:
    kuma.io/mesh: default
spec:
  targetRef:
    kind: Dataplane
    labels:
      app: storefront
  from:
    - targetRef:
        kind: MeshSubset
        tags:
          kuma.io/service: envoy-edge-edge-0a1b2c3d_envoy-gateway-system_svc_443   # from the command above
      default:
        action: Allow

Keep the Mesh in strict mode (the default for a builtin backend when mode is not set) with networking.outbound.passthrough: false, and remove any mesh-wide allow-all permission (see Default-deny MeshTrafficPermission). With passthrough off, a gateway that still routes to pod IPs fails loudly instead of sending plaintext.

Prove it

Run on a lab cluster with Envoy Gateway v1.9.1, Kuma 2.14.3 and Cilium 1.20.2. The backend was an nginx storefront (Service port 8080, container port 80) in the meshed namespace shop; the Gateway sat in edge-mesh, and requests came from outside the cluster through the gateway's NodePort. Each request carried a fake card number in the query string, so a capture shows at once whether the hop is readable.

1. The proxy pods have a sidecar, the controller does not:

text
$ kubectl -n envoy-gateway-system get pods \
    -o custom-columns=NAME:.metadata.name,CONTAINERS:.spec.containers[*].name,INIT:.spec.initContainers[*].name
NAME                                             CONTAINERS               INIT
envoy-edge-mesh-edge-4cfd1dcc-669b88ddbb-z98h8   envoy,shutdown-manager   kuma-init,kuma-sidecar
envoy-gateway-746bf76dfc-x5l9n                   envoy-gateway            <none>

Kuma 2.14 runs kuma-sidecar as a native sidecar, so it shows in the INIT column.

2. Only the sidecar has a token:

text
$ kubectl -n envoy-gateway-system get pod envoy-edge-mesh-edge-4cfd1dcc-669b88ddbb-z98h8 -o json \
    | jq -r '.spec.containers[], .spec.initContainers[] | "\(.name): \([.volumeMounts[]?.mountPath] | join(" "))"'
envoy: /certs /sds
shutdown-manager:
kuma-init: /tmp
kuma-sidecar: /tmp /var/run/secrets/kubernetes.io/serviceaccount

3. The gateway is a delegated gateway Dataplane:

text
$ kubectl -n envoy-gateway-system get dataplanes \
    -o custom-columns=NAME:.metadata.name,GATEWAY-SERVICE:'.spec.networking.gateway.tags.kuma\.io/service',INBOUND:.spec.networking.inbound
NAME                                             GATEWAY-SERVICE                                              INBOUND
envoy-edge-mesh-edge-4cfd1dcc-669b88ddbb-z98h8   envoy-edge-mesh-edge-4cfd1dcc_envoy-gateway-system_svc_443   <none>

A gateway section and no inbound listeners. Do not grep for type: DELEGATED: it is the default and neither kubectl nor kumactl prints it.

4. The hop into the mesh is TLS. Capture on the node that runs the storefront pod, filtered on the pod IP and container port, while you send a request through the gateway:

bash
kubectl debug node/worker-1 -it --image=nicolaka/netshoot --profile=sysadmin -- \
  tcpdump -i any -nn -X -c 20 'dst host 10.244.2.47 and tcp dst port 80'

With routingType: Service, the first data packet from the gateway pod (10.244.2.19) is a TLS ClientHello whose server name is the Kuma service name of the storefront:

text
IP 10.244.2.19.55028 > 10.244.2.47.80: Flags [P.], ... length 1942: HTTP
	0x0030:  4808 c04a 1603 0107 9101 0007 8d03 0351  H..J...........Q
	...
	0x0090:  0738 0000 002b 0029 0000 2673 746f 7265  .8...+.)..&store
	0x00a0:  6672 6f6e 745f 7368 6f70 5f73 7663 5f38  front_shop_svc_8
	0x00b0:  3038 307b 6d65 7368 3d64 6566 6175 6c74  080{mesh=default

1603 starts a TLS record. The card number appeared in 0 captured lines.

5. What the default does. The same request with each routingType and mesh mode:

routingTypeMesh modePassthroughResultCard number on the wire
Servicestrictoff200no, TLS
Endpointstrictoff503no request sent
Endpointstricton503rejected by the storefront's sidecar
Endpointpermissiveon200yes

The last row, from the capture:

text
IP 10.244.2.19.37684 > 10.244.2.47.80: Flags [P.], ...
GET /?card=4111111111111111 HTTP/1.1
x-forwarded-for: 172.18.0.1

The gateway works and the site loads, and the request crosses the node in plain text. Only strict mode turns this into a visible failure.

6. Only the gateway can call the backend:

text
$ kubectl -n app exec deploy/frontend -c app -- \
    curl -s -o /dev/null -w '%{http_code}\n' http://storefront.shop.svc.cluster.local:8080/
403

7. What breaks when a piece is missing. Each piece was removed in turn:

  • Pod label only, namespace unlabelled: the proxy pod starts with no kuma-sidecar (the injector webhook is never called).
  • No token for the sidecar: the sidecar exits with Error: dataplane token is invalid, in Kubernetes you must mount a serviceAccount token, ... could not read file /var/run/secrets/kubernetes.io/serviceaccount/token.
  • No xDS exclusion, passthrough off: the new proxy pod never becomes ready (envoy not ready, kuma-sidecar ready), and Envoy logs gRPC config: initial fetch timed out and DeltaAggregatedResources gRPC config stream to xds_cluster closed since 39s ago: 14, upstream connect error. The old pod keeps serving, so a rollout hides this until the old pod is gone.

Mistakes people make

Keeping the default routingType

Endpoint is the Envoy Gateway default. With a Kuma sidecar it means passthrough, which means no mTLS. Set routingType: Service on the EnvoyProxy, and check any BackendTrafficPolicy that overrides it per route.

Labelling the whole envoy-gateway-system namespace enabled

That puts the controller in the mesh too, and its connections to the Kubernetes API and the proxies go through a sidecar they do not need. Label the namespace disabled and the proxy pods enabled through the EnvoyProxy resource.

Labelling the pods and not the namespace

Kuma's docs say the pod label takes precedence, and it does, but only in a namespace that has the label key. Without it, nothing is injected and nothing warns you.

Turning the token back on for the whole pod

Setting automountServiceAccountToken: true fixes the sidecar and also hands an API token to the internet-facing envoy container. Mount it into the sidecar only.

Forgetting the xDS port

With passthrough off in the mesh, the proxies cannot reach the controller on 18000 and stop receiving configuration. Exclude it from interception (and 18002 if you use Wasm extensions).

Permissive mode "until the gateway works"

The gateway works in permissive mode even when it is sending plaintext. That is why it must be tested in strict mode.

Letting any namespace attach routes

A Gateway with from: All lets any team publish any service on your public hostname. Use a namespace selector.

Checklist

  • The EnvoyProxy sets routingType: Service.
  • The proxy namespace is labelled kuma.io/sidecar-injection: disabled; proxy pods carry kuma.io/sidecar-injection: enabled and kuma.io/gateway: enabled; the controller is not injected.
  • A projected token is mounted into kuma-sidecar only, through a ContainerPatch.
  • Port 18000 is excluded from outbound interception on proxy pods.
  • The mesh runs strict mTLS with passthrough: false and no mesh-wide allow-all.
  • Backends allow the gateway's kuma.io/service in a MeshTrafficPermission, and nothing broader.
  • The Gateway limits route attachment with a namespace selector.
  • A capture on a backend node shows TLS from the gateway, not plain HTTP.

Two good proxies with one wrong default between them is still one plaintext hop. `routingType: Service` is the line that joins them.

H2-CSPE

Learn it on a live range

Service mesh and gateways, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Dome

Want it run for you?

The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.

See the Dome