Kubernetes networking with Cilium

Egress isolation for CI build pods, the secure way

A CI build pod runs whatever the pull request says, including the postinstall script of a package published an hour ago. From inside your cluster network, with your node next door and your metadata endpoint one curl away, that script has a lot of options.

The short answer

Put build pods in their own namespace with no service account token. Give them a Cilium policy that allows only DNS for listed names and HTTPS to the Git server, registry and package mirror by FQDN, and a clusterwide egressDeny for the API server, all nodes, link-local metadata and private ranges, which no later allow can undo.

Updated Houssam Hammoudi, CTOTested with Cilium 1.20.2, Hubble 1.20.2, Kubernetes 1.34 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A build runs code you have not reviewed yet: the pull request, its build scripts, and every dependency they download. A malicious dependency or a hostile pull request gets a shell in the build pod. What it can reach from there is set by the network, not by the CI system.

In a default cluster a build pod can reach:

  • the Kubernetes API, with the pod's service account token if one is mounted;
  • the kubelet and other node ports on every node;
  • the cloud metadata service and the node's cloud credentials;
  • databases and internal services in the VPC;
  • any host on the internet, to fetch a second stage or send stolen secrets.

Network policy that only allows is fragile here. CI configurations change often, and one broad allow added for a new build step re-opens everything.

What the docs say

Deny policies take precedence over allow policies, regardless of whether they are a Cilium Network Policy, a Clusterwide Cilium Network Policy or even a Kubernetes Network Policy.

Source: Cilium docs, Deny Policies

The host entity includes the local host. This also includes all containers running in host networking mode on the local host.

Source: Cilium docs, Layer 3 Policies

CIDR rules apply if Cilium cannot map the source or destination to an identity derived from endpoint labels, ie the Special Identities.

Source: Cilium docs, Layer 3 Policies

The Cilium docs give you the pieces: deny rules that win, entities for the API server and nodes, CIDR rules for everything outside the cluster. They do not assemble them for the workload that most needs them. Note the third quote: CIDR rules do not cover nodes or pods, so a CIDR deny for 10.0.0.0/8 does not block the node next door. The entity deny does.

The secure configuration

1. The build namespace and pod. Builds get no cluster credentials.

yaml
apiVersion: v1
kind: Pod
metadata:
  name: build-example
  namespace: ci-builds
  labels:
    ci.example.com/role: build
spec:
  automountServiceAccountToken: false      # no cluster credentials in a build
  enableServiceLinks: false
  restartPolicy: Never
  containers:
    - name: build
      image: registry.example.com/ci/builder@sha256:0000000000000000000000000000000000000000000000000000000000000000
      securityContext:
        allowPrivilegeEscalation: false
        runAsNonRoot: true
        runAsUser: 10001
        capabilities:
          drop: ["ALL"]
        seccompProfile:
          type: RuntimeDefault

The runner controller that creates build pods is a different workload with its own label and its own, narrow API access. Build pods never get that label.

2. The allowlist: DNS for listed names, HTTPS to named services.

yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: build-pods
  namespace: ci-builds
spec:
  endpointSelector:
    matchLabels:
      ci.example.com/role: build
  ingress:
    - {}                                   # default deny ingress, no rules
  egress:
    - toEndpoints:
        - matchLabels:
            k8s:io.kubernetes.pod.namespace: kube-system
            k8s:k8s-app: kube-dns
      toPorts:
        - ports:
            - port: "53"
              protocol: UDP
            - port: "53"
              protocol: TCP
          rules:
            dns:
              - matchPattern: "**.cluster.local"
              - matchName: "git.example.com"
              - matchName: "registry.example.com"
              - matchName: "mirror.example.net"
    - toFQDNs:
        - matchName: "git.example.com"        # source checkout
        - matchName: "registry.example.com"   # base images and push
        - matchName: "mirror.example.net"     # package mirror (pip, npm, go)
      toPorts:
        - ports:
            - port: "443"
              protocol: TCP

Point every package manager at your mirror. A build that needs the public internet directly is a build that can send data to the public internet.

3. The hard denies. These survive any allow a team adds later.

yaml
apiVersion: cilium.io/v2
kind: CiliumClusterwideNetworkPolicy
metadata:
  name: ci-build-hard-deny
spec:
  endpointSelector:
    matchLabels:
      k8s:io.kubernetes.pod.namespace: ci-builds
      ci.example.com/role: build
  enableDefaultDeny:
    egress: false
    ingress: false
  egressDeny:
    - toEntities:
        - kube-apiserver                   # builds never talk to the cluster API
        - host                             # the node: kubelet, node services
        - remote-node                      # other nodes
    - toCIDRSet:
        - cidr: 10.0.0.0/8
        - cidr: 172.16.0.0/12
        - cidr: 192.168.0.0/16
        - cidr: 100.64.0.0/10
        - cidr: 169.254.0.0/16             # cloud metadata and link-local
        - cidr: fc00::/7
        - cidr: fe80::/10

Deny beats allow, so the private-range deny also blocks any allowlisted service that resolves to a private address, such as an in-VPC mirror. Carve that address out of the range with except on the toCIDRSet entry (for example cidr: 10.0.0.0/8 with except: ["10.20.30.40/32"]) rather than removing the range.

If your cluster runs a node-local DNS cache on a link-local address, the 169.254.0.0/16 and host denies block it. Point build pods at the kube-dns Service instead (with dnsPolicy: None and dnsConfig), or narrow the deny.

4. Run builds on separate nodes. A taint and a node pool for CI keep a build that escapes its container away from production workloads:

yaml
# in the build pod spec
tolerations:
  - key: ci.example.com/builds
    operator: Exists
    effect: NoSchedule
nodeSelector:
  ci.example.com/pool: builds

Prove it

Run on a lab cluster. The example names do not resolve, so the lab policy listed github.com for git, registry-1.docker.io for the registry and pypi.org for the mirror. The checks run from a throwaway pod with the build label:

bash
kubectl -n ci-builds run probe --restart=Never \
  --labels=ci.example.com/role=build --image=curlimages/curl:8.14.1 \
  --overrides='{"spec":{"automountServiceAccountToken":false}}' \
  --command -- sleep infinity

1. Allowed services work:

text
$ kubectl -n ci-builds exec probe -- curl -s -m 8 -o /dev/null -w 'registry %{http_code}\n' https://registry-1.docker.io/v2/
registry 401
git 200
mirror 200

401 is the registry asking for a token: it answered.

2. Everything else fails:

text
$ kubectl -n ci-builds exec probe -- curl -sk -m 5 -o /dev/null -w 'api %{http_code}\n' https://kubernetes.default.svc/version
api 000                (curl exit 28)
metadata 000           (curl exit 28)
kubelet 000            (curl exit 28, https://<node address>:10250/healthz)
internet 000           (curl exit 28, https://example.org/)
$ kubectl -n ci-builds exec probe -- nslookup exfil-test.attacker.example
** server can't find exfil-test.attacker.example: REFUSED

The internet line deserves a second look. example.org is not on the DNS list, so no connection was ever tried, yet curl timed out instead of failing to resolve. The next section explains why, and which setting fixes it.

3. The drops are the policy's:

text
$ hubble observe --namespace ci-builds --type policy-verdict --verdict DROPPED --print-policy-names --since 5m
ci-builds/probe:57618 (ID:8897) <> 172.18.0.3:6443 (kube-apiserver) policy-verdict:L3-Only EGRESS DENIED BY ci-build-hard-deny (CiliumClusterwideNetworkPolicy) (TCP Flags: SYN)
ci-builds/probe:53050 (ID:8897) <> 169.254.169.254:80 (ID:16777218) policy-verdict:L3-Only EGRESS DENIED BY ci-build-hard-deny (CiliumClusterwideNetworkPolicy), deny-cloud-metadata (CiliumClusterwideNetworkPolicy) (TCP Flags: SYN)
ci-builds/probe:34370 (ID:8897) <> 172.18.0.4:10250 (host) policy-verdict:L3-Only EGRESS DENIED BY ci-build-hard-deny (CiliumClusterwideNetworkPolicy) (TCP Flags: SYN)

The metadata drop names two policies: this page's deny and the cluster-wide metadata block. The DNS decisions are in hubble observe --namespace ci-builds --protocol dns. Keep the first query as a dashboard: drops from real builds show attempts you should look at.

Use a Hubble CLI of the same version as Hubble Relay. The lab first ran CLI 1.19.4 against Relay 1.20.2. The CLI printed a warning, then showed every verdict as policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED with no policy names. The flows themselves were fine (-o json had the direction and egress_denied_by). The CLI in the Cilium agent image matches: kubectl -n kube-system exec ds/cilium -c cilium-agent -- hubble version.

4. No token in the build:

text
$ kubectl -n ci-builds exec probe -- ls /var/run/secrets/kubernetes.io/serviceaccount/
ls: /var/run/secrets/kubernetes.io/serviceaccount/: No such file or directory

5. Denied lookups and Alpine images. Cilium's default reject code is refused. The musl C library, used by Alpine-based images such as the curl image here, ignores that answer and keeps asking until its own timeout:

text
dnsRejectResponseCode: refused     1 curl, exit 6 (could not resolve) after 5s
dnsRejectResponseCode: nameError   5 curls, exit 6, all within 1s

Hubble showed the same query sent again every 2.5 seconds under refused. The musl source says so directly:

c
/* Only accept positive or negative responses;
 * retry immediately on server failure, and ignore
 * all other codes such as refusal. */

Source: musl, src/network/res_msend.c

For build pods that run many short commands, five seconds per blocked lookup adds up, and it looks like a network problem instead of a policy decision. Set dnsProxy.dnsRejectResponseCode: nameError so blocked names fail at once with NXDOMAIN.

Mistakes people make

Relying on allow-only policy

CI configuration changes weekly. The first toEntities: [world] added to fix a failing build re-opens everything. The deny policy is what stays true.

Denying 10.0.0.0/8 and thinking nodes are covered

CIDR rules do not apply to nodes and pods by default. Deny host, remote-node and kube-apiserver as entities.

Letting builds reach the internet "just for dependencies"

That is exactly the channel malicious dependencies use. Give builds a mirror, and let only the mirror reach the internet.

Mounting the service account token

A build with a token can talk to the API server if anything allows it, now or later. Set automountServiceAccountToken: false on build pods and on the namespace's default service account.

One policy for runner and builds

The runner controller needs the API; builds do not. Separate them by label so the deny policy cannot be loosened for one without the other.

Checklist

  • Build pods run in a dedicated namespace with automountServiceAccountToken: false.
  • Build pods have a label the runner controller does not share.
  • DNS rules list only Git, registry, mirror and **.cluster.local.
  • toFQDNs allows only those services on 443.
  • A clusterwide egressDeny blocks kube-apiserver, host, remote-node, private ranges and link-local.
  • Build pods have no ingress rules.
  • Builds run on a dedicated, tainted node pool.
  • A probe pod reaches the registry and fails on the API server, metadata, kubelet and the internet.
  • Hubble drops from build pods are watched.

You cannot review every dependency before it runs. You can decide what it can reach when it does.

H2-CSDE

Learn it on a live range

Building without Docker, in DevSecOps and Supply Chain: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on kubernetes networking with cilium

Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.

All kubernetes networking with cilium guides