Kubernetes networking with Cilium
Cilium policy audit mode to enforcement, the secure way
Audit mode sounds like a safe place to test a new network policy. Turn it on for the whole agent, and every policy in the cluster, including the one that blocks the metadata endpoint, quietly stops being enforced.
The short answer
Use Policy Audit Mode per endpoint for the namespace you are rolling out, not daemon-wide on a cluster that already enforces policies. Apply the policy, watch AUDIT verdicts in Hubble until a full business cycle shows none, then disable audit on those endpoints. Re-apply per-endpoint audit after any agent restart, because it resets.
On this page
What goes wrong
Cilium's Policy Audit Mode allows all traffic and logs what policies would have dropped. It is the right tool to find the connections a new policy forgets. It has two scopes, and each has a trap:
- Daemon-wide (
policyAuditMode: true). Every endpoint on every node is in audit mode. No policy is enforced anywhere: not the new one, and not the ones that already protect the cluster. A namespace rollout becomes a cluster-wide outage of network security that lasts until someone turns it off. - Per endpoint (
cilium-dbg endpoint config <id> PolicyAuditMode=Enabled). Only the chosen pods are in audit. But the setting does not survive the agent: when the Cilium pod restarts (an upgrade, a node reboot, a crash), those endpoints go back to the daemon setting and the new policy starts dropping in the middle of your audit.
There is a third, quieter gap: audit mode covers layers 3 and 4. L7 rules (HTTP, DNS, Kafka) are outside it.
What the docs say
When Policy Audit Mode is enabled, no network policy is enforced so this setting is not recommended for production deployment.
Source: Cilium docs, Creating Policies from Verdicts
This approach is meant to be temporary. Restarting Cilium pod will reset the Policy Audit Mode to match the daemon’s configuration.
Source: Cilium docs, Creating Policies from Verdicts
Policy Audit Mode supports auditing network policies implemented at networks layers 3 and 4.
Source: Cilium docs, Creating Policies from Verdicts
The guide is written for a demo cluster with no other policies, where daemon-wide audit costs nothing. It does not say that on a real cluster the same switch turns off the policies you already rely on, and it gives no exit criteria for leaving audit mode.
The secure configuration
1. Decide the scope.
| Situation | Audit scope |
|---|---|
| New cluster, no policies enforced yet | Daemon-wide is acceptable for the first rollout |
| Any policy already enforced (metadata block, default deny elsewhere) | Per endpoint, for the pods under rollout only |
2. Per-endpoint audit for one namespace. This script follows the commands in the Cilium guide, for every endpoint of a namespace:
#!/usr/bin/env bash
# endpoint-audit.sh <namespace> <Enabled|Disabled>
# Sets Policy Audit Mode on every Cilium endpoint of one namespace, through the
# Cilium agent on each pod's node. Temporary by design: an agent restart resets
# endpoints to the daemon setting. Re-run after any Cilium restart or upgrade.
set -euo pipefail
ns=${1:?namespace}
mode=${2:?Enabled or Disabled}
case "$mode" in Enabled|Disabled) ;; *) echo "mode must be Enabled or Disabled" >&2; exit 2 ;; esac
cilium_ns=${CILIUM_NAMESPACE:-kube-system}
kubectl -n "$ns" get ciliumendpoints \
-o jsonpath='{range .items[*]}{.metadata.name}{" "}{.status.id}{"\n"}{end}' |
while read -r pod id; do
node=$(kubectl -n "$ns" get pod "$pod" -o jsonpath='{.spec.nodeName}')
agent=$(kubectl -n "$cilium_ns" get pod -l k8s-app=cilium \
--field-selector spec.nodeName="$node" -o jsonpath='{.items[0].metadata.name}')
kubectl -n "$cilium_ns" exec "$agent" -c cilium-agent -- \
cilium-dbg endpoint config "$id" "PolicyAuditMode=$mode" >/dev/null
now=$(kubectl -n "$cilium_ns" exec "$agent" -c cilium-agent -- \
cilium-dbg endpoint get "$id" -o jsonpath='{[*].spec.options.PolicyAuditMode}')
echo "$ns/$pod endpoint=$id node=$node PolicyAuditMode=$now"
doneThe policy under rollout, in this example: only the frontend may call the backend.
# backend-ingress.yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: backend-ingress
namespace: app
spec:
endpointSelector:
matchLabels:
app.kubernetes.io/name: backend
ingress:
- fromEndpoints:
- matchLabels:
app.kubernetes.io/name: frontend
toPorts:
- ports:
- port: "80"
protocol: TCPOrder matters. Enable audit on the endpoints first, then apply the policy:
./endpoint-audit.sh app Enabled
kubectl apply -f backend-ingress.yamlNew pods (a rollout, a scale-up) start with the daemon setting. Freeze deployments in the namespace during the audit, or re-run the script after each one.
3. Daemon-wide audit, only when nothing is enforced yet.
# cilium-values.yaml
policyAuditMode: true # renders policy-audit-mode: "true"; nothing is enforced4. Exit criteria. Leave audit mode when a full business cycle (a day, a
deploy and the nightly jobs) shows no AUDIT verdicts for the namespace, and
each earlier AUDIT verdict became either a new allow rule or a documented
"this should be blocked".
./endpoint-audit.sh app DisabledProve it
Run on a lab cluster, in a namespace audit with three pods: backend
(behind a Service), frontend, and reports, which the policy does not
allow.
1. Only the intended endpoints are in audit:
$ ./endpoint-audit.sh audit Enabled
audit/backend-bc978fb8b-ptssw endpoint=1151 node=lab-worker2 PolicyAuditMode=Enabled
audit/frontend endpoint=3797 node=lab-worker2 PolicyAuditMode=Enabled
audit/reports endpoint=384 node=lab-worker2 PolicyAuditMode=Enabled
$ kubectl -n kube-system get configmap cilium-config -o jsonpath='{.data.policy-audit-mode}{"\n"}'
The empty line is the daemon setting: not set, so nothing else in the cluster is in audit.
2. What the new policy would drop. With the policy applied, both callers still get through:
frontend -> backend: 200
reports -> backend: 200
$ hubble observe --namespace audit --type policy-verdict --verdict AUDIT --since 5m
audit/reports:53582 (ID:64188) -> audit/backend-bc978fb8b-ptssw:80 (ID:33854) policy-verdict:none INGRESS AUDITED (TCP Flags: SYN)That line is the one connection the policy would block: INGRESS at the
backend, policy-verdict:none because no rule matched it. Decide: a missing
rule, or traffic you mean to stop.
Hubble keeps only recent flows in memory (hubble.eventBufferCapacity, 4095
per node by default), so --since 24h returns only what is still cached. To
cover a full business cycle, write the verdicts to a file with the Hubble
exporter, for example hubble.export.static.enabled: true with
allowList: ['{"verdict":["AUDIT"]}'], and read the file on each node.
3. After a Cilium restart, audit is gone, and the policy enforces:
$ kubectl -n kube-system rollout restart daemonset/cilium
$ cilium-dbg endpoint list -o json | jq -r '.[] | ...'
384 Disabled k8s:app.kubernetes.io/name=reports
1151 Disabled k8s:app.kubernetes.io/name=backend
3797 Disabled k8s:app.kubernetes.io/name=frontend
$ kubectl -n audit exec reports -- curl -s -m 4 -o /dev/null -w '%{http_code}\n' http://backend.audit.svc.cluster.local/
000
command terminated with exit code 28Nobody decided to enforce. An agent restart did, in the middle of the audit.
If reports had been a real dependency, this is an outage.
4. New pods start outside the audit. With audit on and the script not re-run, one scale-up later:
backend-bc978fb8b-ngbmq Disabled
backend-bc978fb8b-ptssw Enabled
frontend Enabled
reports EnabledHalf of the backend now enforces and half audits.
5. After enforcement, the verdicts change:
frontend -> backend: 200
reports -> backend: 000
$ hubble observe --namespace audit --type policy-verdict --since 1m --print-policy-names
audit/reports:52860 (ID:64188) <> audit/backend-bc978fb8b-ptssw:80 (ID:33854) policy-verdict:none INGRESS DENIED (TCP Flags: SYN)
audit/frontend:59522 (ID:60899) -> audit/backend-bc978fb8b-ptssw:80 (ID:33854) policy-verdict:L3-L4 INGRESS ALLOWED BY backend-ingress (CiliumNetworkPolicy) (TCP Flags: SYN)DENIED only for the traffic you chose to block, and no AUDITED left.
Mistakes people make
Daemon-wide audit on a cluster that already enforces
It disables every policy, everywhere, including deny policies that protect the metadata endpoint and the API server. Use per-endpoint audit.
Audit covers deny rules too, even for one endpoint. In the lab, a pod under a
cluster-wide deny for 169.254.169.254 (see
blocking the metadata endpoint):
before: 000 (dropped)
probe in audit: 200
audit off: 000Hubble showed the audited request as to-stack FORWARDED. While a pod is in
audit, your cluster-wide deny rules do not protect it. Audit only the pods of
the rollout, and keep the window short.
Forgetting that per-endpoint audit is temporary
An agent restart resets it to the daemon setting. Your policy starts enforcing halfway through the audit, usually during an upgrade window. Re-apply after every restart, or plan the audit between upgrades.
Auditing an L7 rule
Audit mode covers layers 3 and 4 only. Do not count on it to reveal what an HTTP, DNS or Kafka rule would block. Roll out L7 rules separately, starting from an L7 rule that allows everything, and tighten it while you watch the L7 flows in Hubble.
Leaving audit on "until next sprint"
An endpoint in audit mode enforces nothing. Put an end date on the audit and exit criteria in the change ticket.
New pods during the audit
Pods created during the audit start in the daemon mode. A rollout in the middle of the audit creates enforced pods next to audited ones, and confusing results.
Checklist
- The daemon setting
policy-audit-modeisfalseon clusters that enforce any policy. - Per-endpoint audit is enabled for the rollout namespace before the policy is applied.
- Deployments in the namespace are frozen, or the script re-runs after each rollout.
- The audit is re-applied after every Cilium agent restart.
- L7 rules are rolled out separately from L3/L4 audit.
hubble observe --verdict AUDITis empty for a full business cycle before enforcement.- Every earlier AUDIT verdict became a rule or a documented block.
- Audit mode is disabled at the end, and
cilium-dbg endpoint listconfirms it.
Audit mode is a dress rehearsal. Just make sure it is only your namespace on the stage, not the whole cluster.
H2-CSPE
Learn it on a live range
Cluster networking and policy, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on kubernetes networking with cilium
Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.
All kubernetes networking with cilium guides