Kubernetes networking with Cilium

Default-deny network policies generated from Hubble flows, the secure way

Every pod in the cluster can talk to every other pod, and the one person who knew why left in March. Default deny is the fix everyone agrees on and nobody schedules, because nobody knows what will break.

The short answer

Record real traffic with Hubble for long enough to see every job and batch run, reduce the flows to unique source, destination, port and protocol edges, and write one narrow Cilium policy per edge. Add a namespace-wide default deny for ingress and egress, allow DNS through the DNS proxy, run it in audit mode first, then enforce.

Updated Houssam Hammoudi, CTOTested with Cilium 1.20.2, Hubble 1.20.2, Kubernetes 1.34 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Kubernetes allows all traffic until a policy selects a pod. A cluster without network policy is a flat network: a compromised pod can reach every database, every admin port and the cloud metadata service.

Teams know this. The blocker is fear of breaking things, and it leads to two bad shortcuts:

  • Guessing. Policies written from architecture diagrams miss the nightly job, the metrics scraper, the webhook and the one service that still calls an old API. The first outage gets the policy deleted.
  • Allow what you saw, broadly. Someone watches traffic for ten minutes, sees "app talks to many things", and writes toEntities: [cluster]. The policy exists, the dashboard counts it, and it blocks nothing.

Hubble records every flow Cilium sees, with namespace, workload, labels, port and verdict. That is the input you need, if you collect long enough and turn it into narrow rules.

What the docs say

By default, a pod is non-isolated for ingress; all inbound connections are allowed.

Source: Kubernetes docs, Network Policies

When an endpoint is selected by a network policy, it transitions to a default-deny state, where only explicitly allowed traffic is permitted.

Source: Cilium docs, Policy Enforcement Modes

These steps should be repeated for each connection in the cluster to ensure that the network policy allows all of the expected traffic.

Source: Cilium docs, Creating Policies from Verdicts

The Cilium guide walks through one connection at a time with hubble observe --last 1. It does not tell you that Hubble keeps only a ring buffer of recent flows per node (4095 events by default), so a short observation misses anything that runs hourly, nightly or on deploy.

The secure configuration

1. Make DNS visible, then collect flows long enough. Hubble can name an external destination only if the Cilium DNS proxy saw the lookup, and the proxy only sees lookups that a DNS rule covers. Without one, the backend's call to a payment API shows up as a bare IP address. This policy sends DNS through the proxy and enforces nothing:

yaml
# Collection phase only: DNS visibility, no enforcement.
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: dns-visibility
  namespace: app
spec:
  endpointSelector: {}
  enableDefaultDeny:        # without this, selecting every pod with an egress
    egress: false           # rule would put all of them into default deny
    ingress: false
  egress:
    - toEndpoints:
        - matchLabels:
            io.kubernetes.pod.namespace: kube-system
            k8s-app: kube-dns
      toPorts:
        - ports:
            - port: "53"
              protocol: UDP
            - port: "53"
              protocol: TCP
          rules:
            dns:
              - matchPattern: "*"

Then stream flows for the namespace to a file through at least one full business cycle (a day, plus a deploy, plus the nightly jobs). Streaming with --follow avoids the ring buffer limit:

bash
# Through Hubble Relay (hubble.relay.enabled=true); one namespace at a time.
hubble observe --namespace app --follow --output jsonpb > flows-app.json

For longer windows, enable the Hubble exporter and read its file instead (see Hubble for network forensics).

2. Reduce flows to edges. Hubble records one connection several times: the policy verdict and the packet leaving on the source node, and the packet arriving on the destination node. Each node knows the workload names of its own pods only. This jq program keeps connection starts (TCP SYN without ACK, and UDP), names each side by the label your policies will select, and prints a connection key so each connection counts once:

bash
cat > edges.jq <<'JQ'
# One line per connection: source, destination, port, protocol, connection key.
# Names come from the app.kubernetes.io/name or k8s-app label (what policies select),
# then the workload, then the pod.
def who($e; $ip; $names):
  if $e.namespace then
    $e.namespace + "/" + (
      ([$e.labels[]? | select(startswith("k8s:app.kubernetes.io/name=") or startswith("k8s:k8s-app="))
        | split("=")[1]][0])
      // $e.workloads[0].name
      // $e.pod_name)
  else ($names[0] // $ip) end;
select(.flow != null and (.flow.is_reply // false) == false)
| .flow as $f
| select(($f.l4.TCP.flags.SYN // false) == true and ($f.l4.TCP.flags.ACK // false) == false
         or $f.l4.UDP != null)
| ($f.l4.TCP // $f.l4.UDP) as $l4
| [
    who($f.source; $f.IP.source; []),
    who($f.destination; $f.IP.destination; ($f.destination_names // [])),
    ($l4.destination_port | tostring),
    (if $f.l4.TCP then "TCP" else "UDP" end),
    "\($f.IP.source):\($l4.source_port)>\($f.IP.destination):\($l4.destination_port)"
  ]
| @tsv
JQ

jq -r -f edges.jq flows-app.json | sort -u | cut -f1-4 | sort | uniq -c | sort -rn > edges-app.tsv

Read every line of edges-app.tsv with the service owner. Each line becomes a rule or a question ("why does backend call that address?"). Map workload names to the pod labels your policies will select:

bash
kubectl -n app get deploy backend -o jsonpath='{.spec.selector.matchLabels}{"\n"}'

3. Write the policies: a namespace default deny, DNS, then one rule per edge.

yaml
# 1. Baseline: every pod in the namespace is isolated in both directions.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny
  namespace: app
spec:
  podSelector: {}
  policyTypes:
    - Ingress
    - Egress
---
# 2. DNS for every pod, through the Cilium DNS proxy (needed for toFQDNs).
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: allow-dns
  namespace: app
spec:
  endpointSelector: {}
  egress:
    - toEndpoints:
        - matchLabels:
            io.kubernetes.pod.namespace: kube-system
            k8s-app: kube-dns
      toPorts:
        - ports:
            - port: "53"
              protocol: UDP
            - port: "53"
              protocol: TCP
          rules:
            dns:
              - matchPattern: "*"
---
# 3. Edge: app/frontend -> app/backend 8080/TCP, allowed on both ends.
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: frontend-to-backend
  namespace: app
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: frontend
  egress:
    - toEndpoints:
        - matchLabels:
            app.kubernetes.io/name: backend
      toPorts:
        - ports:
            - port: "8080"
              protocol: TCP
---
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: backend-from-frontend
  namespace: app
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: backend
  ingress:
    - fromEndpoints:
        - matchLabels:
            app.kubernetes.io/name: frontend
      toPorts:
        - ports:
            - port: "8080"
              protocol: TCP
---
# 4. Edge: app/backend -> api.example.com 443/TCP, by name, not by IP.
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: backend-to-payments-api
  namespace: app
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: backend
  egress:
    - toFQDNs:
        - matchName: api.example.com
      toPorts:
        - ports:
            - port: "443"
              protocol: TCP

fromEndpoints and toEndpoints without a namespace label match pods in the policy's own namespace only. For a caller in another namespace, add io.kubernetes.pod.namespace: <namespace> to the selector.

Delete dns-visibility when you apply these: allow-dns takes over DNS.

4. Apply in audit mode first, then enforce. Follow Cilium policy audit mode to enforcement: apply the policies with audit mode on for the namespace's endpoints, watch for AUDIT verdicts, fix the gaps, then turn audit mode off.

Prove it

Run on a lab namespace dd (in place of app): frontend calls backend:8080 every 3 seconds, and backend calls https://example.com/ (in place of api.example.com) every 5 seconds.

0. The collection finds exactly the edges the app has. 45 seconds of flows with dns-visibility applied, through edges.jq:

text
     57 dd/frontend	kube-system/kube-dns	53	UDP
     32 dd/backend	kube-system/kube-dns	53	UDP
     14 dd/frontend	dd/backend	8080	TCP
      8 dd/backend	example.com	443	TCP

Nothing was dropped during collection. For comparison, the same flows without dns-visibility gave dd/backend 104.20.23.154 443 TCP: an address you cannot write a stable rule for. And the first version of the jq program on this page, which named sides by workload only, listed every DNS edge twice (dd/frontend -> kube-system/coredns-66bc5c9577-ssrvf and dd/frontend-55bc847cf-nxd27 -> kube-system/coredns), because each node knows the names of its own pods only.

1. Every pod in the namespace is under policy in both directions:

text
$ cilium-dbg endpoint list
ENDPOINT   POLICY (ingress)   POLICY (egress)   IDENTITY   LABELS (source:key[=value])
620        Enabled            Enabled           17073      k8s:app.kubernetes.io/name=frontend
668        Enabled            Enabled           10007      k8s:app.kubernetes.io/name=backend

Disabled in either column means no policy selects the pod in that direction. Repeat on every node: the list shows local endpoints only.

2. Nothing expected is dropped, and the allows name their policy:

text
$ hubble observe --namespace dd --verdict DROPPED --since 30s
$ hubble observe --namespace dd --type policy-verdict --since 20s --print-policy-names
dd/frontend-55bc847cf-nxd27:43350 (ID:17073) -> dd/backend-869c88c5f7-vftjr:8080 (ID:10007) policy-verdict:L3-L4 EGRESS ALLOWED BY frontend-to-backend (CiliumNetworkPolicy) (TCP Flags: SYN)
dd/frontend-55bc847cf-nxd27:43350 (ID:17073) -> dd/backend-869c88c5f7-vftjr:8080 (ID:10007) policy-verdict:L3-L4 INGRESS ALLOWED BY backend-from-frontend (CiliumNetworkPolicy) (TCP Flags: SYN)

The first command printed nothing: no drops. Each allowed connection shows both ends, each with the policy that allowed it. Use a Hubble CLI of the same version as Hubble Relay: an older CLI printed these same flows as policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN ALLOWED, without names.

3. The deny works from a pod with no rule:

text
$ kubectl -n dd exec probe -- curl -s -m 3 -o /dev/null -w '%{http_code}\n' http://backend.dd.svc.cluster.local:8080/hostname
000
$ hubble observe --namespace dd --type policy-verdict --verdict DROPPED --print-policy-names --since 1m
dd/probe:60424 (ID:22944) <> dd/backend-869c88c5f7-vftjr:8080 (ID:10007) policy-verdict:none EGRESS DENIED (TCP Flags: SYN)

DENIED with no name, and policy-verdict:none: no rule matched, so the default deny dropped it.

Mistakes people make

Ten minutes of flows

The nightly backup, the weekly report and the job that runs on deploy are not in a short capture. Collect across a full cycle, and keep the edge list in version control so the next person can see why each rule exists.

Writing toEntities: cluster and calling it done

A rule that allows the whole cluster is not default deny. Rules should name workloads by label, namespaces by io.kubernetes.pod.namespace, and external services by FQDN.

Allowing egress without allowing the other side's ingress

With default deny in both directions on both namespaces, a connection needs an egress rule at the caller and an ingress rule at the callee. One side only means a drop.

Forgetting DNS

A namespace-wide egress deny also denies DNS. Allow it to kube-dns explicitly, and route it through the DNS proxy so toFQDNs rules work.

Pinning external services by IP from the flow log

The IP in the flow is whatever the DNS name resolved to that day. Use toFQDNs with the name from destination_names; see FQDN egress allowlists with the Cilium DNS proxy.

Assuming the node is blocked too

Kubernetes lets the pod's own node reach it even when the pod is isolated for ingress, and Cilium allows local host traffic by default. Node-level access needs host policies, a separate decision.

Checklist

  • Flows were collected with --follow or the exporter across a full business cycle.
  • Every edge in the edge list has an owner who confirmed it.
  • Each namespace has a default-deny policy for both Ingress and Egress.
  • DNS egress to kube-dns is allowed and goes through the DNS proxy.
  • Each edge has an egress rule at the caller and an ingress rule at the callee.
  • External destinations are allowed by FQDN, not by IP.
  • No policy uses toEntities: [cluster] or world as a shortcut.
  • Policies ran in audit mode and the AUDIT verdicts were reviewed before enforcement.
  • cilium-dbg endpoint list shows enforcement enabled in both directions for every pod.
  • A pod with no rules cannot reach any service in the namespace.

Hubble already watched your cluster for you. Default deny is mostly the discipline of reading what it saw, line by line, before you turn the key.

H2-CSPE

Learn it on a live range

Cluster networking and policy, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on kubernetes networking with cilium

Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.

All kubernetes networking with cilium guides