Logging sandbox egress with Hubble, the secure way
Someone asks where the sandbox from Tuesday afternoon connected to. Hubble saw every packet, and it remembers the last few thousand flows per node, which on a busy node means roughly the last few minutes.
The short answer
Turn on the Hubble exporter with an allowlist for egress from sandbox namespaces and a field mask, so the agent writes only those flows to a rotated file. Add a DNS rule to the sandbox policy so flows carry destination names, ship the file off the node, and restart sandbox pods after enabling it so each connection and its DNS lookup are logged.
On this page
What goes wrong
Hubble, Cilium's flow observer, sees every connection a sandbox pod makes:
source pod, destination, port, verdict. hubble observe makes that feel
like a flow log. It is not one.
Flows live in a ring buffer. Each agent keeps a fixed number of recent flows in memory, 4095 by default. A node that runs many pods fills that in minutes. Anything older is gone, and a restarted agent starts empty.
IP addresses are not answers. A flow to 203.0.113.40:443 tells you
little a week later. Hubble records the DNS name only when Cilium's DNS proxy
saw that pod's lookup, which happens only if a policy with a DNS rule selects
the pod.
Logging everything is not the fix. Exporting all flows from all nodes fills disks and log budgets, and it records every internal call of every application.
Connections that started before you looked. The exporter writes only
what happens after it starts. With the default monitor aggregation
(medium), Cilium reports a connection when it opens, when a new TCP flag
appears, and about once every 5 seconds while packets flow. A connection
opened before the exporter started still shows up through those periodic
reports, but its first packet and the DNS lookup behind it are not in the
file. When you enable logging for a sandbox, restart its pods.
What the docs say
Number of recent flows for Hubble to cache. Defaults to 4095.
Source: Cilium Helm chart values, v1.20.2 (hubble.eventBufferCapacity)
Hubble Exporter is a feature of cilium-agent that lets you write Hubble flows to a file for later consumption as logs.
Source: Cilium docs, Configuring Hubble exporter
In order to associate domain names with IP addresses, Cilium intercepts DNS responses per-Endpoint using a DNS Proxy.
Source: Cilium docs, Layer 3 Examples (DNS based)
Generate a tracing event for send packets only on every new connection, any time a packet contains TCP flags that have not been previously seen for the packet direction, and on average once per monitor-aggregation-interval (assuming that a packet is seen during the interval).
Source: Cilium docs, Kubernetes configuration (monitor-aggregation, level medium)
The exporter page shows how to filter and mask, but not which filter answers "what did this sandbox connect to", and it does not mention that destination names appear only with a DNS rule in place. No page says what the log shows for a connection that was already open when logging began.
The secure configuration
1. Export only sandbox egress, with the fields an investigation needs. Cilium Helm values:
# cilium-hubble-export-values.yaml
hubble:
enabled: true
export:
static:
enabled: true
filePath: /var/run/cilium/hubble/sandbox-egress.log
# Egress flows from pods in the sandbox namespaces only.
allowList:
- '{"source_pod":["sandbox/"],"traffic_direction":["EGRESS"]}'
- '{"source_pod":["labs/"],"traffic_direction":["EGRESS"]}'
# Keep what answers who, where, when and whether it was allowed; drop the rest.
fieldMask:
- time
- node_name
- source.namespace
- source.pod_name
- source.labels
- destination.namespace
- destination.pod_name
- destination_names
- IP
- l4
- verdict
- drop_reason_desc
- traffic_direction
fileMaxSizeMb: 50
fileMaxBackups: 10
fileCompress: truehelm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
--reset-then-reuse-values -f cilium-hubble-export-values.yaml
kubectl -n kube-system rollout restart daemonset/cilium2. Give sandbox flows their DNS names. The DNS rule makes Cilium's DNS
proxy see each lookup, so later flows carry destination_names. Keep your
existing egress rules; this one only adds DNS visibility:
# cilium-sandbox-dns-visibility.yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: dns-visibility
namespace: sandbox
spec:
endpointSelector: {}
egress:
- toEndpoints:
- matchLabels:
k8s:io.kubernetes.pod.namespace: kube-system
k8s-app: kube-dns
toPorts:
- ports:
- port: "53"
protocol: ANY
rules:
dns:
- matchPattern: "*" # every lookup is proxied and recordedThis policy puts the namespace into default-deny for egress: anything not
allowed by it or another policy is dropped. For sandboxes that is the point.
Apply the same policy in every sandbox namespace (labs too).
3. Ship the file off the node. Run your log collector as a DaemonSet
that reads /var/run/cilium/hubble/sandbox-egress.log* with a read-only
hostPath mount and sends each line to your log store, tagged with the node
name. The file on the node is a buffer, not the archive.
4. Restart sandbox pods after turning it on, so each connection they make, and the DNS lookup before it, happens after the exporter and the DNS rule are in place.
Prove it
Run on a lab cluster, with the namespace sbx in place of sandbox (the
allow list entry became "source_pod":["sbx/"]). The sandbox's own egress
rule allowed one external service by name, 172-18-0-1.sslip.io on port
9000: a small TCP server on the lab host.
The agent ConfigMap after the upgrade:
hubble-export-allowlist: {"source_pod":["sbx/"],"traffic_direction":["EGRESS"]} {"source_pod":["labs/"],"traffic_direction":["EGRESS"]}
hubble-export-fieldmask: time node_name source.namespace source.pod_name source.labels destination.namespace destination.pod_name destination_names IP l4 verdict drop_reason_desc traffic_direction
hubble-export-file-compress: true
hubble-export-file-max-backups: 10
hubble-export-file-max-size-mb: 50
hubble-export-file-path: /var/run/cilium/hubble/sandbox-egress.log1. Live view, egress only:
$ hubble observe --namespace sbx --traffic-direction egress --last 40
sbx/test-3:38455 (ID:5913) <- kube-system/coredns-66bc5c9577-snqbj:53 (ID:1734) dns-response proxy FORWARDED (DNS Answer "104.20.23.154,172.66.147.243" TTL: 15 (Proxy example.com. A))
sbx/test-2:34605 (ID:21593) -> 172.18.0.1:9000 (ID:16777219) to-stack FORWARDED (TCP Flags: ACK, PSH)
sbx/test-2:34605 (ID:21593) <- 172.18.0.1:9000 (ID:16777219) to-endpoint FORWARDED (TCP Flags: ACK)The direction filter keeps the replies of egress connections, so both arrows appear.
2. The exporter writes only sandbox egress. The file on the node, one line per distinct flow, with a count:
cat /var/run/cilium/hubble/sandbox-egress.log | jq -c '.flow | {src: .source.pod_name, ns: .source.namespace,
dst: (.destination.pod_name // .IP.destination), names: .destination_names,
port: (.l4.TCP.destination_port // .l4.UDP.destination_port), tcp: .l4.TCP.flags, verdict, dir: .traffic_direction}' \
| sort | uniq -c | sort -rn 6 {"src":"test-3","ns":"sbx","dst":"172.66.147.243","names":["example.com"],"port":443,"tcp":{"SYN":true},"verdict":"DROPPED","dir":"EGRESS"}
6 {"src":"test-3","ns":"sbx","dst":"104.20.23.154","names":["example.com"],"port":443,"tcp":{"SYN":true},"verdict":"DROPPED","dir":"EGRESS"}
5 {"src":"test-1","ns":"sbx","dst":"172.18.0.1","names":null,"port":9000,"tcp":{"PSH":true,"ACK":true},"verdict":"FORWARDED","dir":"EGRESS"}
3 {"src":"test-2","ns":"sbx","dst":"172.18.0.1","names":["172-18-0-1.sslip.io"],"port":9000,"tcp":{"PSH":true,"ACK":true},"verdict":"FORWARDED","dir":"EGRESS"}
2 {"src":"test-2","ns":"sbx","dst":"172.18.0.1","names":["172-18-0-1.sslip.io"],"port":9000,"tcp":{"SYN":true},"verdict":"FORWARDED","dir":"EGRESS"}
2 {"src":"test-1","ns":"sbx","dst":"172.18.0.1","names":null,"port":9000,"tcp":{"PSH":true,"ACK":true},"verdict":"DROPPED","dir":"EGRESS"}(DNS lines to CoreDNS removed from the listing; every line in the file had
"ns":"sbx" and "dir":"EGRESS".)
Only sbx pods, only egress. The blocked attempt to reach example.com
carries its name too: the DNS rule named it before the policy dropped it.
That line is often the one an investigation needs.
3. What an already-open connection leaves in the log. test-1 opened
its connection before the exporter and the DNS rule; test-2 and test-3
started after. The timeline for port 9000 on that node:
02:04:05.920 test-1 9000 ACK,PSH FORWARDED
02:04:15.974 test-1 9000 ACK,PSH FORWARDED
02:04:20.983 test-1 9000 ACK,PSH DROPPED
02:04:20.983 test-1 9000 ACK,PSH DROPPED
02:04:21.077 test-2 53 REDIRECTED
02:04:21.104 test-2 9000 172-18-0-1.sslip.io SYN FORWARDEDTwo things happened to test-1. Its flows never got a name, because its
DNS lookup happened before the rule. And when the DNS policy switched the
namespace to default deny, its open connection was dropped: the address was
not yet known to the FQDN rule. It came back a second later only because
test-2 looked up the same name, which taught Cilium the address, and TCP
retransmitted. Without that coincidence the connection is simply cut. Both
are reasons to restart sandbox pods after turning this on (step 4 of the
configuration).
4. The log survives the buffer (not run here): generate traffic from a
sandbox pod, then run a busy workload on the same node for an hour.
hubble observe --since 1h no longer lists the sandbox flow; the shipped log
still does.
Mistakes people make
Treating hubble observe as a log
It reads a ring buffer of 4095 flows per node by default. Raising the buffer helps a little; exporting is the fix.
Exporting everything
A cluster-wide export records every internal request and grows fast. Filter to the namespaces and direction you need, and mask fields you do not need.
No DNS rule
Without a DNS rule, the exporter records IP addresses only. Cloud and CDN addresses are shared and change often; a week later the IP explains nothing.
Keeping the log on the node
A compromised node, an agent reinstall or a node replacement removes the file. Ship it as it is written.
Enabling logging on running sandboxes
A connection opened before the change appears only through Cilium's
periodic reports: its first packet and its DNS lookup are not in the log,
and without a proxied lookup it has no destination_names. Restart sandbox
pods after enabling the exporter or the DNS rule.
Checklist
- The Hubble static (or dynamic) exporter is enabled.
- Its allowList selects egress from sandbox namespaces only.
- Its fieldMask keeps time, node, source pod, destination, names, L4, verdict and direction.
- File rotation and compression are set.
- A DNS rule selects sandbox pods so flows carry
destination_names. - A collector ships the exporter file off every node.
- Sandbox pods were restarted after logging was enabled.
- Someone can answer "where did sandbox X connect on Tuesday" from the log store.
Hubble has a perfect memory for the last few minutes. For last Tuesday, write it down somewhere that is not the node.
H2-CTDE
Learn it on a live range
Network visibility, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on sandboxing untrusted code
gVisor, pods with no network, and running other people's code without handing them your cluster.
All sandboxing untrusted code guides