Kubernetes networking with Cilium

API server and webhook traffic under default deny, the secure way

You rolled out default deny, the dashboards stayed green, and ten minutes later nobody could create a pod. Or worse: pods create fine, because the policy webhook that should have rejected half of them now times out and its failurePolicy says Ignore.

The short answer

Under default deny, allow two directions explicitly. Pods that call the API get egress to the Cilium kube-apiserver entity, never a toServices rule for default/kubernetes. Webhook pods get ingress from kube-apiserver on their container port, or from nodes where the platform tunnels it. Then audit every webhook's failurePolicy and timeout.

Updated Houssam Hammoudi, CTOTested with Cilium 1.20.2, Hubble 1.20.2, Kubernetes 1.34 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Two kinds of traffic cross the boundary between workloads and the Kubernetes control plane:

  • Pods calling the API server. Operators, controllers, CI runners and anything with a client library. Under egress default deny they need a rule.
  • The API server calling pods. Admission webhooks (policy engines, mesh injectors, cert-manager) are Services backed by pods. Under ingress default deny, the API server's requests to them need a rule too.

Default deny breaks both, and the second one fails in a dangerous way. The API server waits for the webhook (10 seconds by default), then applies the webhook's failurePolicy:

  • Fail: the API request is rejected. Deployments stall, pods do not start, and in the worst case the tools you would use to fix it are blocked too.
  • Ignore: the API request goes through without the webhook. If that webhook was your admission policy engine, every request now skips your policies, and nothing tells you.

The obvious rules also fail. The API server is not an ordinary Service: the default/kubernetes Service has no pod selector, and on self-managed clusters the API server runs on nodes, which CIDR rules do not match by default.

What the docs say

The special Kubernetes Service default/kubernetes does not use a label selector. It is not recommended to grant access to the Kubernetes API server with a toServices-based policy.

Source: Cilium docs, Layer 3 Policies

The kube-apiserver entity may not work for ingress traffic in some Kubernetes distributions, such as Azure AKS and GCP GKE.

Source: Cilium docs, Layer 3 Policies

If the timeout expires before the webhook responds, the webhook call will be ignored or the API call will be rejected based on the failure policy.

Source: Kubernetes docs, Dynamic Admission Control

By default, CIDR-based selectors do not match in-cluster entities (pods or nodes).

Source: Cilium docs, Layer 3 Policies

The Cilium docs describe the entity and its limits; the Kubernetes docs describe failure policies. Neither says what happens when you combine default deny with a policy webhook set to Ignore: your admission policy turns off without an error.

The secure configuration

1. Egress to the API server by identity.

yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: operator-to-apiserver
  namespace: app
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: app-operator
  egress:
    - toEntities:
        - kube-apiserver          # in-cluster or external API server, by identity
      toPorts:
        - ports:
            - port: "6443"
              protocol: TCP
            - port: "443"
              protocol: TCP

Give this rule only to pods that need the API. Most application pods do not; they should also run with automountServiceAccountToken: false.

2. Ingress from the API server to each webhook. Allow the container port behind the webhook Service (the targetPort), not the Service port.

yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: webhook-from-apiserver
  namespace: policy-system
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: policy-webhook
  ingress:
    - fromEntities:
        - kube-apiserver
      toPorts:
        - ports:
            - port: "9443"        # the container port behind the webhook Service
              protocol: TCP

On platforms where control plane traffic arrives through worker nodes (Cilium names AKS and GKE), the source is a node, not the kube-apiserver identity. Use the node entities instead, which is broader:

yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: webhook-from-nodes
  namespace: policy-system
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: policy-webhook
  ingress:
    - fromEntities:
        - remote-node
        - host
      toPorts:
        - ports:
            - port: "9443"
              protocol: TCP

If you run a self-managed control plane on labelled nodes, you can narrow this with fromNodes and a node-role.kubernetes.io/control-plane selector, which requires nodeSelectorLabels: true in Cilium.

3. Inventory every webhook, with its failure policy and timeout.

bash
cat > webhook-check.jq <<'JQ'
.items[] as $c
| $c.webhooks[]
| [ $c.kind, $c.metadata.name, .name, (.failurePolicy // "Fail"),
    ((.timeoutSeconds // 10) | tostring),
    (if .clientConfig.service then .clientConfig.service.namespace + "/" + .clientConfig.service.name
     else (.clientConfig.url // "?") end) ]
| @tsv
JQ

kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurations -o json \
  | jq -r -f webhook-check.jq

For every line: the webhook's namespace needs the ingress rule above before default deny goes on. Security webhooks (admission policy, image signature verification) should use failurePolicy: Fail with a short timeout, so a broken path is loud, not silent. Exempt kube-system and the webhook's own namespace with a namespaceSelector so a failure cannot lock you out.

Prove it

Run on a kind cluster, where the API server runs on the control-plane node. The webhook is a small lab server in policy-system that rejects any pod whose image name contains unsigned, with failurePolicy: Fail and timeoutSeconds: 5. It covers the namespace wh-app.

1. What default deny does to a webhook without the ingress rule:

text
$ time kubectl -n wh-app run good --image=busybox:1.37 --dry-run=server -o name
Error from server (InternalError): Internal error occurred: failed calling webhook "images.policy.example.com": failed to call webhook: Post "https://policy-webhook.policy-system.svc:443/?timeout=5s": context deadline exceeded
5.13s
$ hubble observe --namespace policy-system --verdict DROPPED --print-policy-names --since 1m
172.18.0.3:32942 (kube-apiserver) <> policy-system/policy-webhook-7884f577f8-55st9:9443 (ID:41229) policy-verdict:none INGRESS DENIED (TCP Flags: SYN)
172.18.0.3:32942 (kube-apiserver) <> policy-system/policy-webhook-7884f577f8-55st9:9443 (ID:41229) Policy denied DROPPED (TCP Flags: SYN)

Every pod create in the namespace now waits 5 seconds and fails. Before the default deny, the same command took 0.13 seconds.

2. With webhook-from-apiserver:

text
$ time kubectl -n wh-app run good --image=busybox:1.37 --dry-run=server -o name
pod/good
0.14s
$ hubble observe --namespace policy-system --type policy-verdict --print-policy-names --since 20s
172.18.0.3:60656 (kube-apiserver) -> policy-system/policy-webhook-7884f577f8-55st9:9443 (ID:41229) policy-verdict:L3-L4 INGRESS ALLOWED BY webhook-from-apiserver (CiliumNetworkPolicy) (TCP Flags: SYN)
$ hubble observe --from-label reserved:kube-apiserver --verdict DROPPED --since 20s

The last command printed nothing: no more drops from the API server.

3. A security webhook still rejects what it should:

text
$ kubectl -n wh-app run bad-probe --image=registry.example.com/unsigned:latest --dry-run=server
Error from server: admission webhook "images.policy.example.com" denied the request: image not signed: registry.example.com/unsigned:latest

And the reason for failurePolicy: Fail. The same request with the webhook unreachable and the policy set to Ignore:

text
$ time kubectl -n wh-app run bad-probe --image=registry.example.com/unsigned:latest --dry-run=server -o name
pod/bad-probe
5.15s

The unsigned image was admitted. The only sign was 5 seconds of delay.

4. The inventory (the lab also runs Kuma, whose webhooks show up too):

text
$ kubectl get validatingwebhookconfigurations,mutatingwebhookconfigurations -o json | jq -r -f webhook-check.jq
ValidatingWebhookConfiguration	kuma-validating-webhook-configuration	validator.kuma-admission.kuma.io	Fail	10	kuma-system/kuma-control-plane
ValidatingWebhookConfiguration	kuma-validating-webhook-configuration	secret.validator.kuma-admission.kuma.io	Ignore	10	kuma-system/kuma-control-plane
ValidatingWebhookConfiguration	lab-image-policy	images.policy.example.com	Fail	5	policy-system/policy-webhook
MutatingWebhookConfiguration	kuma-admission-mutating-webhook-configuration	pods-kuma-injector.kuma.io	Fail	10	kuma-system/kuma-control-plane

(Trimmed to four of nine lines.) Each namespace in the last column needs the ingress rule before default deny goes on there.

5. Operators can reach the API, others cannot. wh-app under default deny, DNS allowed, and operator-to-apiserver for app-operator only:

text
$ kubectl -n wh-app exec deploy/app-operator -- sh -c 'wget -qO- --timeout=5 https://kubernetes.default.svc/healthz --no-check-certificate; echo " rc=$?"'
ok rc=0
$ kubectl -n wh-app exec deploy/frontend -- sh -c '...same...'
 rc=1
$ hubble observe --to-label reserved:kube-apiserver --verdict DROPPED --since 1m
wh-app/frontend-5b59cdf448-g2vmw:58538 (ID:15045) <> 172.18.0.3:6443 (kube-apiserver) Policy denied DROPPED (TCP Flags: SYN)

The pod asked for kubernetes.default.svc:443, and the policy saw 172.18.0.3:6443: Cilium applies egress policy after the Service is translated to the API server's own address and port. That is why the rule lists 6443 as well as 443.

Mistakes people make

Allowing default/kubernetes with toServices

That Service has no selector, and Cilium recommends against it. Use the kube-apiserver entity.

Using ipBlock for the API server in a Kubernetes NetworkPolicy

On self-managed clusters the API server runs on nodes, and Cilium's CIDR selectors do not match nodes unless policyCIDRMatchMode: nodes is set. The rule looks right and allows nothing.

Opening the Service port instead of the pod port

The API server calls the Service, but the policy sees the pod. Allow the container port the Service targets.

Security webhooks with failurePolicy Ignore

With Ignore, a network drop turns your admission policy off. Use Fail, short timeouts, and namespace exemptions for the system namespaces.

Forgetting AKS and GKE tunnel control plane traffic

There the kube-apiserver entity may not match webhook calls. Test the webhook path after rollout; fall back to node entities if it fails.

Checklist

  • Pods that need the API have egress to toEntities: [kube-apiserver] on 6443 and 443.
  • No policy uses toServices for default/kubernetes.
  • Every webhook's pods allow ingress from kube-apiserver (or node entities where required) on the container port.
  • Security webhooks use failurePolicy: Fail with a short timeoutSeconds.
  • Webhooks exempt kube-system and their own namespace with a namespaceSelector.
  • hubble observe shows no drops to or from reserved:kube-apiserver.
  • A server-side dry run returns quickly, and a bad request is still rejected.

Default deny is easy for the traffic you think about. The API server calling your pods is the traffic nobody draws on the diagram, and it holds the keys.

H2-CSPE

Learn it on a live range

Cluster networking and policy, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on kubernetes networking with cilium

Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.

All kubernetes networking with cilium guides