Kubernetes networking with Cilium
Egress isolation for CI build pods, the secure way
A CI build pod runs whatever the pull request says, including the postinstall script of a package published an hour ago. From inside your cluster network, with your node next door and your metadata endpoint one curl away, that script has a lot of options.
The short answer
Put build pods in their own namespace with no service account token. Give them a Cilium policy that allows only DNS for listed names and HTTPS to the Git server, registry and package mirror by FQDN, and a clusterwide egressDeny for the API server, all nodes, link-local metadata and private ranges, which no later allow can undo.
On this page
What goes wrong
A build runs code you have not reviewed yet: the pull request, its build scripts, and every dependency they download. A malicious dependency or a hostile pull request gets a shell in the build pod. What it can reach from there is set by the network, not by the CI system.
In a default cluster a build pod can reach:
- the Kubernetes API, with the pod's service account token if one is mounted;
- the kubelet and other node ports on every node;
- the cloud metadata service and the node's cloud credentials;
- databases and internal services in the VPC;
- any host on the internet, to fetch a second stage or send stolen secrets.
Network policy that only allows is fragile here. CI configurations change often, and one broad allow added for a new build step re-opens everything.
What the docs say
Deny policies take precedence over allow policies, regardless of whether they are a Cilium Network Policy, a Clusterwide Cilium Network Policy or even a Kubernetes Network Policy.
Source: Cilium docs, Deny Policies
The host entity includes the local host. This also includes all containers running in host networking mode on the local host.
Source: Cilium docs, Layer 3 Policies
CIDR rules apply if Cilium cannot map the source or destination to an identity derived from endpoint labels, ie the Special Identities.
Source: Cilium docs, Layer 3 Policies
The Cilium docs give you the pieces: deny rules that win, entities for the
API server and nodes, CIDR rules for everything outside the cluster. They do
not assemble them for the workload that most needs them. Note the third
quote: CIDR rules do not cover nodes or pods, so a CIDR deny for 10.0.0.0/8
does not block the node next door. The entity deny does.
The secure configuration
1. The build namespace and pod. Builds get no cluster credentials.
apiVersion: v1
kind: Pod
metadata:
name: build-example
namespace: ci-builds
labels:
ci.example.com/role: build
spec:
automountServiceAccountToken: false # no cluster credentials in a build
enableServiceLinks: false
restartPolicy: Never
containers:
- name: build
image: registry.example.com/ci/builder@sha256:0000000000000000000000000000000000000000000000000000000000000000
securityContext:
allowPrivilegeEscalation: false
runAsNonRoot: true
runAsUser: 10001
capabilities:
drop: ["ALL"]
seccompProfile:
type: RuntimeDefaultThe runner controller that creates build pods is a different workload with its own label and its own, narrow API access. Build pods never get that label.
2. The allowlist: DNS for listed names, HTTPS to named services.
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: build-pods
namespace: ci-builds
spec:
endpointSelector:
matchLabels:
ci.example.com/role: build
ingress:
- {} # default deny ingress, no rules
egress:
- toEndpoints:
- matchLabels:
k8s:io.kubernetes.pod.namespace: kube-system
k8s:k8s-app: kube-dns
toPorts:
- ports:
- port: "53"
protocol: UDP
- port: "53"
protocol: TCP
rules:
dns:
- matchPattern: "**.cluster.local"
- matchName: "git.example.com"
- matchName: "registry.example.com"
- matchName: "mirror.example.net"
- toFQDNs:
- matchName: "git.example.com" # source checkout
- matchName: "registry.example.com" # base images and push
- matchName: "mirror.example.net" # package mirror (pip, npm, go)
toPorts:
- ports:
- port: "443"
protocol: TCPPoint every package manager at your mirror. A build that needs the public internet directly is a build that can send data to the public internet.
3. The hard denies. These survive any allow a team adds later.
apiVersion: cilium.io/v2
kind: CiliumClusterwideNetworkPolicy
metadata:
name: ci-build-hard-deny
spec:
endpointSelector:
matchLabels:
k8s:io.kubernetes.pod.namespace: ci-builds
ci.example.com/role: build
enableDefaultDeny:
egress: false
ingress: false
egressDeny:
- toEntities:
- kube-apiserver # builds never talk to the cluster API
- host # the node: kubelet, node services
- remote-node # other nodes
- toCIDRSet:
- cidr: 10.0.0.0/8
- cidr: 172.16.0.0/12
- cidr: 192.168.0.0/16
- cidr: 100.64.0.0/10
- cidr: 169.254.0.0/16 # cloud metadata and link-local
- cidr: fc00::/7
- cidr: fe80::/10Deny beats allow, so the private-range deny also blocks any allowlisted
service that resolves to a private address, such as an in-VPC mirror. Carve
that address out of the range with except on the toCIDRSet entry (for
example cidr: 10.0.0.0/8 with except: ["10.20.30.40/32"]) rather than
removing the range.
If your cluster runs a node-local DNS cache on a link-local address, the
169.254.0.0/16 and host denies block it. Point build pods at the
kube-dns Service instead (with dnsPolicy: None and dnsConfig), or narrow
the deny.
4. Run builds on separate nodes. A taint and a node pool for CI keep a build that escapes its container away from production workloads:
# in the build pod spec
tolerations:
- key: ci.example.com/builds
operator: Exists
effect: NoSchedule
nodeSelector:
ci.example.com/pool: buildsProve it
Run on a lab cluster. The example names do not resolve, so the lab policy
listed github.com for git, registry-1.docker.io for the registry and
pypi.org for the mirror. The checks run from a throwaway pod with the build
label:
kubectl -n ci-builds run probe --restart=Never \
--labels=ci.example.com/role=build --image=curlimages/curl:8.14.1 \
--overrides='{"spec":{"automountServiceAccountToken":false}}' \
--command -- sleep infinity1. Allowed services work:
$ kubectl -n ci-builds exec probe -- curl -s -m 8 -o /dev/null -w 'registry %{http_code}\n' https://registry-1.docker.io/v2/
registry 401
git 200
mirror 200401 is the registry asking for a token: it answered.
2. Everything else fails:
$ kubectl -n ci-builds exec probe -- curl -sk -m 5 -o /dev/null -w 'api %{http_code}\n' https://kubernetes.default.svc/version
api 000 (curl exit 28)
metadata 000 (curl exit 28)
kubelet 000 (curl exit 28, https://<node address>:10250/healthz)
internet 000 (curl exit 28, https://example.org/)
$ kubectl -n ci-builds exec probe -- nslookup exfil-test.attacker.example
** server can't find exfil-test.attacker.example: REFUSEDThe internet line deserves a second look. example.org is not on the DNS
list, so no connection was ever tried, yet curl timed out instead of failing
to resolve. The next section explains why, and which setting fixes it.
3. The drops are the policy's:
$ hubble observe --namespace ci-builds --type policy-verdict --verdict DROPPED --print-policy-names --since 5m
ci-builds/probe:57618 (ID:8897) <> 172.18.0.3:6443 (kube-apiserver) policy-verdict:L3-Only EGRESS DENIED BY ci-build-hard-deny (CiliumClusterwideNetworkPolicy) (TCP Flags: SYN)
ci-builds/probe:53050 (ID:8897) <> 169.254.169.254:80 (ID:16777218) policy-verdict:L3-Only EGRESS DENIED BY ci-build-hard-deny (CiliumClusterwideNetworkPolicy), deny-cloud-metadata (CiliumClusterwideNetworkPolicy) (TCP Flags: SYN)
ci-builds/probe:34370 (ID:8897) <> 172.18.0.4:10250 (host) policy-verdict:L3-Only EGRESS DENIED BY ci-build-hard-deny (CiliumClusterwideNetworkPolicy) (TCP Flags: SYN)The metadata drop names two policies: this page's deny and the cluster-wide
metadata block. The DNS
decisions are in hubble observe --namespace ci-builds --protocol dns. Keep
the first query as a dashboard: drops from real builds show attempts you
should look at.
Use a Hubble CLI of the same version as Hubble Relay. The lab first ran CLI
1.19.4 against Relay 1.20.2. The CLI printed a warning, then showed every
verdict as policy-verdict:none TRAFFIC_DIRECTION_UNKNOWN DENIED with no
policy names. The flows themselves were fine (-o json had the direction and
egress_denied_by). The CLI in the Cilium agent image matches:
kubectl -n kube-system exec ds/cilium -c cilium-agent -- hubble version.
4. No token in the build:
$ kubectl -n ci-builds exec probe -- ls /var/run/secrets/kubernetes.io/serviceaccount/
ls: /var/run/secrets/kubernetes.io/serviceaccount/: No such file or directory5. Denied lookups and Alpine images. Cilium's default reject code is
refused. The musl C library, used by Alpine-based images such as the curl
image here, ignores that answer and keeps asking until its own timeout:
dnsRejectResponseCode: refused 1 curl, exit 6 (could not resolve) after 5s
dnsRejectResponseCode: nameError 5 curls, exit 6, all within 1sHubble showed the same query sent again every 2.5 seconds under refused.
The musl source says so directly:
/* Only accept positive or negative responses;
* retry immediately on server failure, and ignore
* all other codes such as refusal. */Source: musl, src/network/res_msend.c
For build pods that run many short commands, five seconds per blocked lookup
adds up, and it looks like a network problem instead of a policy decision.
Set dnsProxy.dnsRejectResponseCode: nameError so blocked names fail at once
with NXDOMAIN.
Mistakes people make
Relying on allow-only policy
CI configuration changes weekly. The first toEntities: [world] added to fix a
failing build re-opens everything. The deny policy is what stays true.
Denying 10.0.0.0/8 and thinking nodes are covered
CIDR rules do not apply to nodes and pods by default. Deny host,
remote-node and kube-apiserver as entities.
Letting builds reach the internet "just for dependencies"
That is exactly the channel malicious dependencies use. Give builds a mirror, and let only the mirror reach the internet.
Mounting the service account token
A build with a token can talk to the API server if anything allows it, now or
later. Set automountServiceAccountToken: false on build pods and on the
namespace's default service account.
One policy for runner and builds
The runner controller needs the API; builds do not. Separate them by label so the deny policy cannot be loosened for one without the other.
Checklist
- Build pods run in a dedicated namespace with
automountServiceAccountToken: false. - Build pods have a label the runner controller does not share.
- DNS rules list only Git, registry, mirror and
**.cluster.local. toFQDNsallows only those services on 443.- A clusterwide
egressDenyblockskube-apiserver,host,remote-node, private ranges and link-local. - Build pods have no ingress rules.
- Builds run on a dedicated, tainted node pool.
- A probe pod reaches the registry and fails on the API server, metadata, kubelet and the internet.
- Hubble drops from build pods are watched.
You cannot review every dependency before it runs. You can decide what it can reach when it does.
H2-CSDE
Learn it on a live range
Building without Docker, in DevSecOps and Supply Chain: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on kubernetes networking with cilium
Default-deny network policy, transparent encryption and egress control with Cilium and Hubble.
All kubernetes networking with cilium guides