Giving learners root in a sandbox safely, the secure way
Your lab promises "full root access", and the first learner took it literally: they tried to read the node's disk, scan the cluster network and fork until the node gave up. To be fair, the syllabus did say "explore".
The short answer
Give each learner a pod where root is not host root: a gVisor RuntimeClass, or hostUsers false on nodes with user namespace support. No ServiceAccount token, no host access, no Service links. Deny all network traffic except to that learner's own targets, cap CPU, memory, disk and processes, and end every session with activeDeadlineSeconds.
On this page
What goes wrong
A Linux lab where the learner is not root teaches half the syllabus. So lab
platforms hand out root, usually as UID 0 in a normal container on runc.
That root shares the host kernel. It holds the default container
capabilities, including NET_RAW, and it can use every system call the
default seccomp profile allows. One kernel bug reachable from a container
turns a curious learner into the owner of the node, and every other
learner's session with it.
The rest of the damage needs no bug at all:
- The pod's ServiceAccount token is mounted, so
curlto the API server works as that account. - Service environment variables list every Service in the namespace.
- The network is flat: the learner can reach other learners' pods, internal Services and the cloud metadata endpoint.
- A fork bomb or a
ddinto/tmpexhausts the node's processes or disk. - The session never ends, so a crypto miner started on day one is still running on day thirty.
What the docs say
This behavior does not present a security concern because root inside a Pod with user namespaces actually refers to the user inside the container, that is never mapped to a privileged user on the host.
Source: Kubernetes docs, User Namespaces
An application in a gVisor sandbox is permitted to do most things a standard container can do: for example, applications can read and write files mapped within the container, make network connections, etc.
Source: gVisor docs, Security Model
gVisor similarly relies on the host resource mechanisms (cgroups) for defense against resource exhaustion and denial of service attacks.
Source: gVisor docs, Security Model
Both mechanisms take root away from the host. Neither limits what root can reach on the network, how much it can consume, or how long it lives. Those are separate controls, and the docs for each assume you add them.
The secure configuration
Pick one containment. gVisor (runtimeClassName: gvisor, see
gVisor for untrusted workloads)
gives root a separate kernel. User namespaces (hostUsers: false, stable
since Kubernetes 1.36) map root to an unprivileged host UID on runc; they need
idmap mount support on the node's filesystems (Linux 6.3 or later in
practice), containerd 2.0 or later and runc 1.2 or later. The example uses
gVisor.
1. The namespace, its quota and its defaults.
# k8s-labs-namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
name: labs
labels:
sandbox.example.com/untrusted: "true" # the gVisor admission policy applies
pod-security.kubernetes.io/enforce: baseline # root allowed; host access and privileged are not
pod-security.kubernetes.io/enforce-version: v1.37
---
apiVersion: v1
kind: ResourceQuota
metadata:
name: labs
namespace: labs
spec:
hard:
pods: "200"
requests.cpu: "50"
limits.memory: 100Gi
requests.ephemeral-storage: 200Gi
services: "0" # learners get no Services
---
apiVersion: v1
kind: LimitRange
metadata:
name: labs
namespace: labs
spec:
limits:
- type: Container
default:
cpu: "1"
memory: 512Mi
ephemeral-storage: 2Gi
defaultRequest:
cpu: 100m
memory: 256Mi
ephemeral-storage: 1Gi
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny
namespace: labs
spec:
podSelector: {}
policyTypes: ["Ingress", "Egress"] # nothing in, nothing out, unless a learner policy allows it2. One pod per learner session. The lab controller creates it from this template:
# k8s-learner-pod.yaml
apiVersion: v1
kind: Pod
metadata:
name: lab-4821
namespace: labs
labels:
app: lab
learner: "4821"
spec:
runtimeClassName: gvisor # root talks to the Sentry, not the host kernel
automountServiceAccountToken: false # no API credential in the pod
enableServiceLinks: false # no Service addresses in the environment
activeDeadlineSeconds: 7200 # the session ends after two hours, whatever happens
terminationGracePeriodSeconds: 5
hostname: lab
securityContext:
runAsUser: 0 # the learner is root, on purpose
seccompProfile:
type: RuntimeDefault
containers:
- name: shell
image: ghcr.io/example/lab-linux@sha256:0000000000000000000000000000000000000000000000000000000000000000
command: ["sleep", "infinity"] # learners attach through the exec gateway
securityContext:
capabilities:
drop: ["NET_RAW"] # no raw sockets on runc; runsc already removes it unless --net-raw is set
resources:
requests:
cpu: 200m
memory: 256Mi
ephemeral-storage: 1Gi
limits:
cpu: "1"
memory: 512Mi
ephemeral-storage: 2Gi # a full /tmp evicts this pod, not the node3. Each learner reaches only their own targets. The controller creates one policy per session:
# k8s-learner-netpol.yaml
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: learner-4821
namespace: labs
spec:
podSelector:
matchLabels:
learner: "4821"
policyTypes: ["Ingress", "Egress"]
ingress:
- from:
- podSelector:
matchLabels:
learner: "4821" # the learner's shell and their target pods
egress:
- to:
- podSelector:
matchLabels:
learner: "4821"No DNS, no internet, no other learner, no metadata endpoint. If a lab needs the internet, route it through a proxy with an allowlist, never directly.
4. Cap processes on the sandbox nodes. The kubelet's per-pod PID limit bounds the host processes each pod can create, so one session cannot exhaust the node's PID space. It counts host processes only. Under gVisor the Sentry implements the application's threading model, so processes inside the sandbox are not host processes and this limit does not count them; there, gVisor relies on the host cgroups (the pod's memory and CPU limits) to contain resource exhaustion. On Talos:
# talos-lab-kubelet.yaml
# Sandbox worker pool only.
apiVersion: v1alpha1
kind: KubeletConfig
config:
podPidsLimit: 512 # maximum PIDs in any pod on this node5. Access through a gateway, not SSH. Learners reach the shell with
kubectl exec through a gateway that allows exec into their own pod only
(see Exec-only access to sandboxes).
Nothing listens inside the pod.
Prove it
Run on a lab cluster (Kubernetes 1.34, Cilium 1.20.2, gVisor release-20260921)
with the namespace, pod and policy above; the learner image placeholder was
replaced by busybox:1.37, and a target pod labeled learner: "4821" served
HTTP on 8080. Every command runs inside the learner's shell. Real output:
1. Root, but inside gVisor:
id; dmesg | head -2uid=0(root) gid=0(root) groups=0(root),10(wheel)
[ 0.000000] Starting gVisor...
[ 0.307497] Adversarially training Redcode AI...2. No credentials and no service map:
ls /var/run/secrets/kubernetes.io/serviceaccount; env | grep -c _SERVICE_HOSTls: /var/run/secrets/kubernetes.io/serviceaccount: No such file or directory
1The 1 is KUBERNETES_SERVICE_HOST, which Kubernetes sets in every pod.
Before enableServiceLinks: false, there is one line per Service in the
namespace.
3. The network ends at the learner's own pods:
wget -qO- -T 3 http://169.254.169.254/ ; echo "metadata exit $?"
wget -qO- -T 3 https://kubernetes.default.svc/ ; echo "api exit $?"
wget -qO- -T 3 http://<target-pod-ip>:8080/hostname ; echo "own target exit $?"wget: download timed out
metadata exit 1
wget: bad address 'kubernetes.default.svc'
api exit 1
target-4821
own target exit 04. Resource limits hold:
dd if=/dev/zero of=/tmp/fill bs=1M count=3000
kubectl -n labs describe pod lab-4821 # from outside, afterwardswaiting on pid 15: ... urpc method "containerManager.WaitPID" failed: EOF
command terminated with exit code 128
Last State: Terminated
Reason: OOMKilled
Exit Code: 128The session stopped, the container restarted, and the node stayed Ready.
On the lab the writes counted as memory, so the 512Mi memory limit fired
first. gVisor's docs say only that the disk-backed overlay is "typically
enabled by default"; when it is, the same writes count against the 2Gi
ephemeral-storage limit and the kubelet evicts the pod instead. Either way
the learner fills their own sandbox, never the node: check describe to see
which limit your setup hits.
Mistakes people make
Root on runc, "because it is only a lab"
A lab is a room full of people whose job is to try things. Root on the host kernel is one kernel bug away from the node. Use gVisor or user namespaces.
Dropping the ServiceAccount token but keeping Services
enableServiceLinks defaults to true, and every Service in the namespace
becomes a set of environment variables. Turn it off and create no Services in the
learner namespace.
One namespace, one policy, many learners
A policy that allows traffic "within the namespace" lets learners attack each other. Scope policies by a per-learner label.
No deadline
A session without activeDeadlineSeconds lives until someone notices it.
Set the deadline on the pod, so the cluster enforces it even if the lab
controller is down.
Exempting the lab namespace from Pod Security
Root does not require the privileged profile. baseline allows UID 0 and
still blocks privileged containers, host namespaces and hostPath.
Checklist
- Learner pods run with
runtimeClassName: gvisororhostUsers: false. - The namespace enforces the
baselinePod Security level and is not exempt. automountServiceAccountToken: falseandenableServiceLinks: falseare set.- The namespace has default-deny ingress and egress; each learner has a policy scoped to their own label.
- CPU, memory and ephemeral storage have limits; a LimitRange sets defaults.
podPidsLimitis set on sandbox nodes.- Every session pod has
activeDeadlineSeconds. - Learners reach pods through an exec gateway; nothing in the pod listens for them.
NET_RAWis dropped, andrunscdoes not run with--net-raw.
Give learners root, a kernel that is not yours, and a room with no doors. They will learn everything the lab has to teach, and nothing about your nodes.
H2-CSPE
Learn it on a live range
Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on sandboxing untrusted code
gVisor, pods with no network, and running other people's code without handing them your cluster.
All sandboxing untrusted code guides