Sandboxing untrusted code

Logging sandbox egress with Hubble, the secure way

Someone asks where the sandbox from Tuesday afternoon connected to. Hubble saw every packet, and it remembers the last few thousand flows per node, which on a busy node means roughly the last few minutes.

The short answer

Turn on the Hubble exporter with an allowlist for egress from sandbox namespaces and a field mask, so the agent writes only those flows to a rotated file. Add a DNS rule to the sandbox policy so flows carry destination names, ship the file off the node, and restart sandbox pods after enabling it so each connection and its DNS lookup are logged.

Updated Houssam Hammoudi, CTOTested with Cilium 1.20.2, Hubble 1.20.2, Kubernetes 1.34 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Hubble, Cilium's flow observer, sees every connection a sandbox pod makes: source pod, destination, port, verdict. hubble observe makes that feel like a flow log. It is not one.

Flows live in a ring buffer. Each agent keeps a fixed number of recent flows in memory, 4095 by default. A node that runs many pods fills that in minutes. Anything older is gone, and a restarted agent starts empty.

IP addresses are not answers. A flow to 203.0.113.40:443 tells you little a week later. Hubble records the DNS name only when Cilium's DNS proxy saw that pod's lookup, which happens only if a policy with a DNS rule selects the pod.

Logging everything is not the fix. Exporting all flows from all nodes fills disks and log budgets, and it records every internal call of every application.

Connections that started before you looked. The exporter writes only what happens after it starts. With the default monitor aggregation (medium), Cilium reports a connection when it opens, when a new TCP flag appears, and about once every 5 seconds while packets flow. A connection opened before the exporter started still shows up through those periodic reports, but its first packet and the DNS lookup behind it are not in the file. When you enable logging for a sandbox, restart its pods.

What the docs say

Number of recent flows for Hubble to cache. Defaults to 4095.

Source: Cilium Helm chart values, v1.20.2 (hubble.eventBufferCapacity)

Hubble Exporter is a feature of cilium-agent that lets you write Hubble flows to a file for later consumption as logs.

Source: Cilium docs, Configuring Hubble exporter

In order to associate domain names with IP addresses, Cilium intercepts DNS responses per-Endpoint using a DNS Proxy.

Source: Cilium docs, Layer 3 Examples (DNS based)

Generate a tracing event for send packets only on every new connection, any time a packet contains TCP flags that have not been previously seen for the packet direction, and on average once per monitor-aggregation-interval (assuming that a packet is seen during the interval).

Source: Cilium docs, Kubernetes configuration (monitor-aggregation, level medium)

The exporter page shows how to filter and mask, but not which filter answers "what did this sandbox connect to", and it does not mention that destination names appear only with a DNS rule in place. No page says what the log shows for a connection that was already open when logging began.

The secure configuration

1. Export only sandbox egress, with the fields an investigation needs. Cilium Helm values:

yaml
# cilium-hubble-export-values.yaml
hubble:
  enabled: true
  export:
    static:
      enabled: true
      filePath: /var/run/cilium/hubble/sandbox-egress.log
      # Egress flows from pods in the sandbox namespaces only.
      allowList:
        - '{"source_pod":["sandbox/"],"traffic_direction":["EGRESS"]}'
        - '{"source_pod":["labs/"],"traffic_direction":["EGRESS"]}'
      # Keep what answers who, where, when and whether it was allowed; drop the rest.
      fieldMask:
        - time
        - node_name
        - source.namespace
        - source.pod_name
        - source.labels
        - destination.namespace
        - destination.pod_name
        - destination_names
        - IP
        - l4
        - verdict
        - drop_reason_desc
        - traffic_direction
      fileMaxSizeMb: 50
      fileMaxBackups: 10
      fileCompress: true
bash
helm upgrade cilium cilium/cilium --version 1.20.2 --namespace kube-system \
  --reset-then-reuse-values -f cilium-hubble-export-values.yaml
kubectl -n kube-system rollout restart daemonset/cilium

2. Give sandbox flows their DNS names. The DNS rule makes Cilium's DNS proxy see each lookup, so later flows carry destination_names. Keep your existing egress rules; this one only adds DNS visibility:

yaml
# cilium-sandbox-dns-visibility.yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: dns-visibility
  namespace: sandbox
spec:
  endpointSelector: {}
  egress:
    - toEndpoints:
        - matchLabels:
            k8s:io.kubernetes.pod.namespace: kube-system
            k8s-app: kube-dns
      toPorts:
        - ports:
            - port: "53"
              protocol: ANY
          rules:
            dns:
              - matchPattern: "*"     # every lookup is proxied and recorded

This policy puts the namespace into default-deny for egress: anything not allowed by it or another policy is dropped. For sandboxes that is the point. Apply the same policy in every sandbox namespace (labs too).

3. Ship the file off the node. Run your log collector as a DaemonSet that reads /var/run/cilium/hubble/sandbox-egress.log* with a read-only hostPath mount and sends each line to your log store, tagged with the node name. The file on the node is a buffer, not the archive.

4. Restart sandbox pods after turning it on, so each connection they make, and the DNS lookup before it, happens after the exporter and the DNS rule are in place.

Prove it

Run on a lab cluster, with the namespace sbx in place of sandbox (the allow list entry became "source_pod":["sbx/"]). The sandbox's own egress rule allowed one external service by name, 172-18-0-1.sslip.io on port 9000: a small TCP server on the lab host.

The agent ConfigMap after the upgrade:

text
hubble-export-allowlist: {"source_pod":["sbx/"],"traffic_direction":["EGRESS"]} {"source_pod":["labs/"],"traffic_direction":["EGRESS"]}
hubble-export-fieldmask: time node_name source.namespace source.pod_name source.labels destination.namespace destination.pod_name destination_names IP l4 verdict drop_reason_desc traffic_direction
hubble-export-file-compress: true
hubble-export-file-max-backups: 10
hubble-export-file-max-size-mb: 50
hubble-export-file-path: /var/run/cilium/hubble/sandbox-egress.log

1. Live view, egress only:

text
$ hubble observe --namespace sbx --traffic-direction egress --last 40
sbx/test-3:38455 (ID:5913) <- kube-system/coredns-66bc5c9577-snqbj:53 (ID:1734) dns-response proxy FORWARDED (DNS Answer "104.20.23.154,172.66.147.243" TTL: 15 (Proxy example.com. A))
sbx/test-2:34605 (ID:21593) -> 172.18.0.1:9000 (ID:16777219) to-stack FORWARDED (TCP Flags: ACK, PSH)
sbx/test-2:34605 (ID:21593) <- 172.18.0.1:9000 (ID:16777219) to-endpoint FORWARDED (TCP Flags: ACK)

The direction filter keeps the replies of egress connections, so both arrows appear.

2. The exporter writes only sandbox egress. The file on the node, one line per distinct flow, with a count:

bash
cat /var/run/cilium/hubble/sandbox-egress.log | jq -c '.flow | {src: .source.pod_name, ns: .source.namespace,
  dst: (.destination.pod_name // .IP.destination), names: .destination_names,
  port: (.l4.TCP.destination_port // .l4.UDP.destination_port), tcp: .l4.TCP.flags, verdict, dir: .traffic_direction}' \
  | sort | uniq -c | sort -rn
text
      6 {"src":"test-3","ns":"sbx","dst":"172.66.147.243","names":["example.com"],"port":443,"tcp":{"SYN":true},"verdict":"DROPPED","dir":"EGRESS"}
      6 {"src":"test-3","ns":"sbx","dst":"104.20.23.154","names":["example.com"],"port":443,"tcp":{"SYN":true},"verdict":"DROPPED","dir":"EGRESS"}
      5 {"src":"test-1","ns":"sbx","dst":"172.18.0.1","names":null,"port":9000,"tcp":{"PSH":true,"ACK":true},"verdict":"FORWARDED","dir":"EGRESS"}
      3 {"src":"test-2","ns":"sbx","dst":"172.18.0.1","names":["172-18-0-1.sslip.io"],"port":9000,"tcp":{"PSH":true,"ACK":true},"verdict":"FORWARDED","dir":"EGRESS"}
      2 {"src":"test-2","ns":"sbx","dst":"172.18.0.1","names":["172-18-0-1.sslip.io"],"port":9000,"tcp":{"SYN":true},"verdict":"FORWARDED","dir":"EGRESS"}
      2 {"src":"test-1","ns":"sbx","dst":"172.18.0.1","names":null,"port":9000,"tcp":{"PSH":true,"ACK":true},"verdict":"DROPPED","dir":"EGRESS"}

(DNS lines to CoreDNS removed from the listing; every line in the file had "ns":"sbx" and "dir":"EGRESS".)

Only sbx pods, only egress. The blocked attempt to reach example.com carries its name too: the DNS rule named it before the policy dropped it. That line is often the one an investigation needs.

3. What an already-open connection leaves in the log. test-1 opened its connection before the exporter and the DNS rule; test-2 and test-3 started after. The timeline for port 9000 on that node:

text
02:04:05.920  test-1  9000                         ACK,PSH  FORWARDED
02:04:15.974  test-1  9000                         ACK,PSH  FORWARDED
02:04:20.983  test-1  9000                         ACK,PSH  DROPPED
02:04:20.983  test-1  9000                         ACK,PSH  DROPPED
02:04:21.077  test-2  53                                    REDIRECTED
02:04:21.104  test-2  9000  172-18-0-1.sslip.io    SYN      FORWARDED

Two things happened to test-1. Its flows never got a name, because its DNS lookup happened before the rule. And when the DNS policy switched the namespace to default deny, its open connection was dropped: the address was not yet known to the FQDN rule. It came back a second later only because test-2 looked up the same name, which taught Cilium the address, and TCP retransmitted. Without that coincidence the connection is simply cut. Both are reasons to restart sandbox pods after turning this on (step 4 of the configuration).

4. The log survives the buffer (not run here): generate traffic from a sandbox pod, then run a busy workload on the same node for an hour. hubble observe --since 1h no longer lists the sandbox flow; the shipped log still does.

Mistakes people make

Treating hubble observe as a log

It reads a ring buffer of 4095 flows per node by default. Raising the buffer helps a little; exporting is the fix.

Exporting everything

A cluster-wide export records every internal request and grows fast. Filter to the namespaces and direction you need, and mask fields you do not need.

No DNS rule

Without a DNS rule, the exporter records IP addresses only. Cloud and CDN addresses are shared and change often; a week later the IP explains nothing.

Keeping the log on the node

A compromised node, an agent reinstall or a node replacement removes the file. Ship it as it is written.

Enabling logging on running sandboxes

A connection opened before the change appears only through Cilium's periodic reports: its first packet and its DNS lookup are not in the log, and without a proxied lookup it has no destination_names. Restart sandbox pods after enabling the exporter or the DNS rule.

Checklist

  • The Hubble static (or dynamic) exporter is enabled.
  • Its allowList selects egress from sandbox namespaces only.
  • Its fieldMask keeps time, node, source pod, destination, names, L4, verdict and direction.
  • File rotation and compression are set.
  • A DNS rule selects sandbox pods so flows carry destination_names.
  • A collector ships the exporter file off every node.
  • Sandbox pods were restarted after logging was enabled.
  • Someone can answer "where did sandbox X connect on Tuesday" from the log store.

Hubble has a perfect memory for the last few minutes. For last Tuesday, write it down somewhere that is not the node.

H2-CTDE

Learn it on a live range

Network visibility, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on sandboxing untrusted code

gVisor, pods with no network, and running other people's code without handing them your cluster.

All sandboxing untrusted code guides