Detecting and stopping sandbox abuse, the secure way
Your sandboxes are isolated, limited and logged, and one of them has used exactly one CPU core, its limit, for nineteen hours. The logs say "compiling". The destination port says 3333.
The short answer
Alert on three signals per sandbox pod: policy drops, which mean scanning or probing; new connections to the internet on mining and mail ports; and CPU held at its limit for a long time. When one fires, save the evidence first, then revoke the user's access and block the pod's network; delete the pod only once the evidence is stored. Test the alert rules first.
On this page
What goes wrong
A sandbox that runs anyone's code will be used for things you did not plan: mining cryptocurrency, scanning the internet or your own network, sending spam, probing the metadata service and the API server, or simply holding a foothold. Good isolation turns most of that into failed attempts. Nobody looks at failed attempts unless something counts them.
Typical gaps:
- Drops are logged but not counted, so a sandbox that tries ten thousand addresses looks the same as one that tried one.
- CPU limits hide mining. A miner at its limit does not hurt the node, so no node alert fires, and it runs for days on your bill.
- Responders delete the pod first. The pod's metrics series disappear a minute later, the logs go with it, and the evidence that would explain the incident, or clear the user, is gone.
- Access is revoked last. The user starts a new sandbox while you are deleting the old one.
What the docs say
To limit metrics cardinality hubble will remove data series bound to specific pod after one minute from pod deletion.
Source: Cilium docs, Monitoring & Metrics
By default, dropped flows are counted if and only if the drop reason is Policy denied.
Source: Cilium docs, Monitoring & Metrics (flows-to-world)
it is implementation defined as to whether the change will take effect for that existing connection or not.
Source: Kubernetes docs, Network Policies
The Cilium page explains the metrics but not which ones detect abuse, and it buries the one-minute cleanup of pod series in a note. Kubernetes says a policy change that denies a connection may leave an attacker's open connection alive, which matters when you quarantine instead of deleting.
The secure configuration
1. Hubble metrics with the sandbox pod as a label. Cilium Helm values:
# cilium-hubble-metrics-values.yaml
hubble:
metrics:
enabled:
# Drops per source pod: scanning and probing show up here.
- "drop:labelsContext=source_namespace,source_pod"
# New TCP connections to the internet, with the destination port.
- "flows-to-world:port;syn-only;labelsContext=source_namespace,source_pod"Only sandbox namespaces should have pod-level labels if the cluster is
large; the series count grows with pods. By default flows-to-world counts
forwarded flows and flows dropped with the reason Policy denied, so its
verdict label tells an attempt from a connection that went through.
2. Alert rules. For the Prometheus Operator:
# k8s-sandbox-abuse-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: sandbox-abuse
namespace: monitoring
spec:
groups:
- name: sandbox-abuse
rules:
- alert: SandboxPolicyDropsHigh
# A sandbox hitting its network policy over and over: scanning or probing.
expr: sum by (source_namespace, source_pod) (rate(hubble_drop_total{source_namespace=~"sandbox|labs"}[5m])) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "{{ $labels.source_namespace }}/{{ $labels.source_pod }} is dropping {{ $value | humanize }} packets/s at policy"
- alert: SandboxMiningOrMailPort
# New connections to the internet on common mining pool and mail ports.
expr: sum by (source_namespace, source_pod, port) (increase(hubble_flows_to_world_total{source_namespace=~"sandbox|labs", port=~"25|465|587|3333|4444|5555|7777|14444"}[10m])) > 0
labels:
severity: critical
annotations:
summary: "{{ $labels.source_namespace }}/{{ $labels.source_pod }} opened connections to port {{ $labels.port }}"
- alert: SandboxCPUAtLimit
# Pinned at the CPU limit for an hour: the classic miner profile.
# The pod-level row (container="", image=""): gVisor pods have no per-container rows.
expr: |
sum by (namespace, pod) (rate(container_cpu_usage_seconds_total{namespace=~"sandbox|labs", container="", image=""}[10m]))
/ sum by (namespace, pod) (kube_pod_container_resource_limits{namespace=~"sandbox|labs", resource="cpu"})
> 0.9
for: 1h
labels:
severity: warning
annotations:
summary: "{{ $labels.namespace }}/{{ $labels.pod }} has used over 90% of its CPU limit for an hour"Route critical to a person, not only a channel. SandboxCPUAtLimit reads
the pod-level row of container_cpu_usage_seconds_total on purpose. Under
gVisor the application runs inside the Sentry, and cAdvisor has no row for
the application container at all, only the pod and the pause container
(Prove it, step 3). A rule that filters on container!="" never fires for a
gVisor sandbox. The flow log from
Logging sandbox egress with Hubble
answers the next question: where did it connect?
3. A response script: evidence first, access second, the pod last.
# sandbox-quarantine.sh
# Usage: bash sandbox-quarantine.sh <namespace> <pod> <session-label-value>
# Needs a Hubble Relay connection: HUBBLE_SERVER, and HUBBLE_TLS=true,
# HUBBLE_TLS_CA_CERT_FILES and HUBBLE_TLS_SERVER_NAME when Relay uses TLS.
# It never deletes the pod: it isolates it, and a person deletes it once the
# evidence is stored somewhere safe.
set -euo pipefail
ns="$1"; pod="$2"; session="$3"
case_dir="case-$(date -u +%Y%m%dT%H%M%SZ)-${ns}-${pod}"
mkdir "$case_dir" # fails if it exists: never write into an old case
umask 077
fail() { echo "$1; nothing was changed. Partial evidence in $case_dir" >&2; exit 1; }
# 1. Evidence, while the pod, its flows and its access objects still exist.
kubectl -n "$ns" get pod "$pod" -o yaml > "$case_dir/pod.yaml"
kubectl -n "$ns" describe pod "$pod" > "$case_dir/pod-describe.txt"
kubectl -n "$ns" logs "$pod" --all-containers --timestamps > "$case_dir/logs.txt" || true
kubectl -n "$ns" logs "$pod" --all-containers --timestamps --previous > "$case_dir/logs-previous.txt" 2>/dev/null || true
hubble observe --namespace "$ns" --pod "$pod" --since 24h -o jsonpb > "$case_dir/flows.json" \
|| fail "hubble observe failed"
kubectl -n "$ns" get events --field-selector "involvedObject.name=$pod" -o yaml > "$case_dir/events.yaml"
kubectl -n "$ns" get role,rolebinding -l "sandbox.example.com/session=$session" -o yaml > "$case_dir/access.yaml"
# Every capture that must have content, has content. Logs may be empty.
for f in pod.yaml pod-describe.txt flows.json events.yaml access.yaml; do
[ -s "$case_dir/$f" ] || fail "$f is empty"
done
(cd "$case_dir" && sha256sum -- * > SHA256SUMS)
chmod -R a-w "$case_dir" # the evidence is read-only from here on
# 2. Access: remove the user's exec binding (its copy is in access.yaml).
kubectl -n "$ns" delete role,rolebinding -l "sandbox.example.com/session=$session"
# 3. Network: label the pod for the clusterwide deny policy (drops new and open connections).
kubectl -n "$ns" label pod "$pod" network.example.com/none=true --overwrite
echo "evidence in $case_dir (check with: cd $case_dir && sha256sum -c SHA256SUMS)"
echo "the pod is isolated, not deleted. After the evidence is copied off this machine:"
echo " kubectl -n $ns delete pod $pod --now"The network.example.com/none label is matched by the clusterwide deny
policy in Pods with no network at all. Then
disable the user in your identity provider; the sandbox was only one of their
doors.
Prove it
Run on a lab cluster, with the namespace sbx standing in for sandbox,
plus promtool unit tests for the rules.
1. The metrics carry pod labels. After a sandbox pod tried port 3333 and port 25 on an outside address, and another probed 30 ports:
$ curl -s localhost:9965/metrics | grep -E '^hubble_(drop|flows_to_world)_total' | grep sbx
hubble_drop_total{protocol="TCP",reason="POLICY_DENIED",source_namespace="sbx",source_pod="test-2"} 24
hubble_drop_total{protocol="TCP",reason="POLICY_DENIED",source_namespace="sbx",source_pod="test-3"} 60
hubble_flows_to_world_total{port="25",protocol="TCP",source_namespace="sbx",source_pod="test-2",verdict="DROPPED"} 6
hubble_flows_to_world_total{port="3333",protocol="TCP",source_namespace="sbx",source_pod="test-2",verdict="DROPPED"} 18
hubble_flows_to_world_total{port="8001",protocol="TCP",source_namespace="sbx",source_pod="test-3",verdict="DROPPED"} 2The port-3333 attempts are counted although the policy dropped every SYN,
with verdict="DROPPED". That is what the mining alert keys on.
2. The rules fire on those shapes. The unit test
(secure-tests/detecting-stopping-sandbox-abuse/rules-test.yaml) feeds a pod
dropping 5 packets per second, a pod opening connections to port 3333, a runc
pod at 98% of its limit, a gVisor pod over its limit, and a quiet pod:
Checking rules.yaml
SUCCESS: 3 rules found
SUCCESS3. What cAdvisor reports under gVisor. Two pods burning CPU with a
200m limit, one on runc and one on the gvisor RuntimeClass:
container_cpu_usage_seconds_total{container="",...,pod="burn-runc"}
container_cpu_usage_seconds_total{container="",image="registry.k8s.io/pause:3.10",...,pod="burn-runc"}
container_cpu_usage_seconds_total{container="burn",image="docker.io/library/busybox:1.37",...,pod="burn-runc"}
container_cpu_usage_seconds_total{container="",image="",name="",namespace="cpu-test",pod="burn-gvisor"}
container_cpu_usage_seconds_total{container="",image="registry.k8s.io/pause:3.10",...,pod="burn-gvisor"}No container="burn" row for the gVisor pod. The first version of this
page's rule filtered on container!="", and a promtool test with these
gVisor series returned got:[]: the alert could never fire for a sandbox.
The rule now uses the pod row (container="", image=""), which both runtimes
have.
The gVisor pod also used 0.3 cores against its 0.2-core limit. Its pod
cgroup had cpu.max 30000 100000: the container limit plus the RuntimeClass
overhead, enforced for the whole sandbox rather than per container. The
alert divides by the container limits, so a gVisor pod at its cgroup limit
shows 150%, which still fires.
4. The quarantine script. The first version of this page ran
hubble observe ... || true and deleted the pod at the end. With no Relay
settings, hubble failed, the script went on, and the pod was deleted with
an empty flows.json. The script above never deletes the pod, and changes
nothing until every capture has content. Without Hubble settings:
$ bash sandbox-quarantine.sh sbx test-5 ghi
rpc error: code = Unavailable desc = connection error: desc = "error reading server preface: EOF"
hubble observe failed; nothing was changed. Partial evidence in case-20260925T060025Z-sbx-test-5
$ kubectl -n sbx get pod test-5 --no-headers; kubectl -n sbx get rolebinding -l sandbox.example.com/session=ghi --no-headers
test-5 1/1 Running 0 25s
session-ghi Role/session-ghi 23sWith HUBBLE_SERVER, HUBBLE_TLS=true, HUBBLE_TLS_CA_CERT_FILES and
HUBBLE_TLS_SERVER_NAME set:
$ bash sandbox-quarantine.sh sbx test-5 ghi
role.rbac.authorization.k8s.io "session-ghi" deleted from sbx namespace
rolebinding.rbac.authorization.k8s.io "session-ghi" deleted from sbx namespace
pod/test-5 labeled
evidence in case-20260925T060028Z-sbx-test-5 (check with: cd case-20260925T060028Z-sbx-test-5 && sha256sum -c SHA256SUMS)
the pod is isolated, not deleted. After the evidence is copied off this machine:
kubectl -n sbx delete pod test-5 --now
$ ls -l case-20260925T060028Z-sbx-test-5
-r-------- 550 SHA256SUMS
-r-------- 901 access.yaml
-r-------- 2952 events.yaml
-r-------- 349349 flows.json
-r-------- 0 logs-previous.txt
-r-------- 0 logs.txt
-r-------- 2315 pod-describe.txt
-r-------- 3319 pod.yaml
$ cd case-20260925T060028Z-sbx-test-5 && sha256sum -c SHA256SUMS
access.yaml: OK
events.yaml: OK
flows.json: OK
...
pod.yaml: OKThe pod kept running, labelled and cut off, with the session's Role and
RoleBinding saved in access.yaml before they were removed. The checksums
are written inside the folder, so they still verify after the folder is
copied elsewhere. Read-only files stop accidents, not root: copy the folder
to write-once storage before anyone deletes the pod.
5. The quarantine label cuts open connections too. A pod with an open
connection, labelled network.example.com/none=true:
sbx/test-2:34605 (ID:17418) <> 172.18.0.1:9000 (ID:16777219) policy-verdict:all EGRESS DENIED BY no-network (CiliumClusterwideNetworkPolicy) (TCP Flags: ACK, PSH)
sbx/test-2:34605 (ID:17418) <> 172.18.0.1:9000 (ID:16777219) Policy denied by denylist DROPPED (TCP Flags: ACK, PSH)
$ kubectl -n sbx exec test-2 -- nc -w 3 172-18-0-1.sslip.io 9000
nc: bad address '172-18-0-1.sslip.io'Data on the open connection is dropped, and new ones cannot even resolve a
name. The sockets stay ESTABLISHED on both ends until the pod is deleted,
which a person does once the evidence is stored.
Mistakes people make
Deleting first, asking later
The pod's logs vanish with it, and Hubble drops its metric series a minute after deletion. Save the pod spec, logs, events, flows and access objects, store them off the machine, then delete.
Quarantine by policy alone
With Cilium, the deny label dropped data on an open connection too (Prove it, step 5), but the process keeps running and the sockets stay open. Delete the pod once the evidence is stored off the machine.
Alerting on node CPU
A miner held to its limit never raises node CPU enough to page anyone. Compare each sandbox pod's usage with its own limit.
Counting only the internet
flows-to-world counts only flows whose destination has the
reserved:world identity. Probes against the API server and other pods never
reach that identity, so only the drop metric counts them.
Revoking the pod but not the person
The user can start another sandbox a minute later. Remove their bindings, then their identity-provider access, then clean up.
Checklist
- Hubble
dropandflows-to-worldmetrics carrysource_namespaceandsource_podfor sandbox namespaces. - Alerts exist for policy drops, mining and mail ports, and CPU at limit.
- The rules pass
promtool check rulesand a unit test with synthetic series. - Critical alerts page a person.
- The response saves pod spec, logs, events, flows and the user's bindings before changing anything.
- It stops without changes if any capture is empty.
- Evidence files are hashed, read-only, and copied to write-once storage.
- The pod is isolated by label, and deleted by a person only after the evidence is stored.
- The user is disabled in the identity provider.
Isolation makes abuse fail; detection makes it visible; a good runbook makes it boring. Aim for boring.
H2-CTDE
Learn it on a live range
Alerting and incident response, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.
Start freeH2 Security services
Want it done with your team?
Our engineers set it up with you, test it the way this page does, and leave it documented.
See our services