Sandboxing untrusted code

One pod per code execution, the secure way

The "run this snippet" button was a hit, so the team kept a warm pool of interpreter pods to make it fast. The second user's code found the first user's variables still in memory, and a file called results.csv.

The short answer

Give every execution a new pod that is deleted afterwards: gVisor with no network, no ServiceAccount token, non-root, read-only root filesystem, input from a per-run ConfigMap, output through logs, a short deadline and hard limits. Let the executor create only such pods, enforced by a ValidatingAdmissionPolicy, in a namespace that holds nothing else.

Updated Houssam Hammoudi, CTOTested with Kubernetes 1.34.0 (kind), gVisor release-20260921 (runsc-nonet)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A feature that runs user code, a playground, an autograder, an AI agent's code tool, is remote code execution on purpose. The design decides how much of your system each execution touches.

The fast designs share things:

  • Warm pools and long-lived interpreters. Execution B runs in the same process or pod as execution A, and finds A's files, environment and memory.
  • The service's own pod. Code runs inside the API server that received it, next to its database credentials.
  • A broad executor role. The service that creates execution pods can create any pod in its namespace. If it is compromised, or tricked by input it passes into a pod spec, it creates a privileged pod with a hostPath mount and the node is gone.
  • Leftovers. Finished pods and their ConfigMaps stay around, with user code and output inside, until someone cleans up.

What the docs say

Permission to create workloads (either Pods, or workload resources that manage Pods) in a namespace implicitly grants access to many other resources in that namespace, such as Secrets, ConfigMaps, and PersistentVolumes that can be mounted in Pods.

Source: Kubernetes docs, Role Based Access Control Good Practices

Safely running untrusted code, such as when running third-party/user-provided code, or for software forensics.

Source: gVisor docs, Production guide (use cases)

For failed Pods, the API objects remain in the cluster's API until a human or controller process explicitly removes them.

Source: Kubernetes docs, Pod Lifecycle

RBAC can say "create pods", not "create pods that look like this", so the shape of the pod has to be enforced by admission. And finished pods are only garbage-collected once the cluster passes the terminated-pod-gc-threshold of the controller manager, so cleanup is the executor's job.

The secure configuration

1. A namespace that holds executions and nothing else, restricted by Pod Security, labeled for the gVisor policy, with no Secrets.

yaml
# k8s-exec-namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: code-exec
  labels:
    sandbox.example.com/untrusted: "true"
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/enforce-version: v1.37
---
apiVersion: v1
kind: ServiceAccount
metadata:
  name: default
  namespace: code-exec
automountServiceAccountToken: false
---
# The executor runs elsewhere (namespace "executor") and may only manage pods,
# ConfigMaps and logs here.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: executor
  namespace: code-exec
rules:
  - apiGroups: [""]
    resources: ["pods", "configmaps"]
    verbs: ["create", "get", "delete"]
  - apiGroups: [""]
    resources: ["pods/log"]
    verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: executor
  namespace: code-exec
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: executor
subjects:
  - kind: ServiceAccount
    name: executor
    namespace: executor

2. The pod the executor creates for each run. Code arrives in a ConfigMap made for this run; the result is the log.

yaml
# k8s-exec-pod.yaml
apiVersion: v1
kind: Pod
metadata:
  name: exec-9d41c2
  namespace: code-exec
  labels:
    app: code-exec
    network.example.com/none: "true"     # the clusterwide deny policy applies too
spec:
  runtimeClassName: gvisor-nonet         # gVisor with --network=none: only loopback
  restartPolicy: Never
  activeDeadlineSeconds: 30
  terminationGracePeriodSeconds: 0
  automountServiceAccountToken: false
  enableServiceLinks: false
  securityContext:
    runAsNonRoot: true
    runAsUser: 65534
    runAsGroup: 65534
    seccompProfile:
      type: RuntimeDefault
  containers:
    - name: run
      image: ghcr.io/example/python-runner@sha256:0000000000000000000000000000000000000000000000000000000000000000
      command: ["python3", "-I", "/code/main.py"]   # -I: ignore env vars and user site dirs
      securityContext:
        allowPrivilegeEscalation: false
        readOnlyRootFilesystem: true
        capabilities:
          drop: ["ALL"]
      resources:
        requests:
          cpu: 250m
          memory: 128Mi
          ephemeral-storage: 64Mi
        limits:
          cpu: "1"
          memory: 256Mi
          ephemeral-storage: 128Mi
      volumeMounts:
        - name: code
          mountPath: /code
          readOnly: true
        - name: tmp
          mountPath: /tmp
  volumes:
    - name: code
      configMap:
        name: exec-9d41c2                  # created for this run, deleted with the pod
    - name: tmp
      emptyDir:
        sizeLimit: 64Mi

3. Enforce that shape for everything the executor creates. RBAC lets the executor create pods; this policy decides which pods.

yaml
# k8s-exec-policy.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: code-exec-pod-shape
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
      - apiGroups: [""]
        apiVersions: ["v1"]
        operations: ["CREATE"]
        resources: ["pods"]
  validations:
    - expression: "has(object.spec.runtimeClassName) && object.spec.runtimeClassName == 'gvisor-nonet'"
      message: "execution pods must use the gvisor-nonet RuntimeClass"
    - expression: "has(object.spec.automountServiceAccountToken) && object.spec.automountServiceAccountToken == false"
      message: "execution pods must not mount a ServiceAccount token"
    - expression: "has(object.spec.activeDeadlineSeconds) && object.spec.activeDeadlineSeconds <= 60"
      message: "execution pods must set activeDeadlineSeconds of 60 or less"
    - expression: "has(object.spec.restartPolicy) && object.spec.restartPolicy == 'Never'"
      message: "execution pods must not restart"
    - expression: "!has(object.spec.volumes) || object.spec.volumes.all(v, has(v.configMap) || has(v.emptyDir))"
      message: "execution pods may mount only ConfigMaps and emptyDir"
    - expression: "!has(object.spec.initContainers) || size(object.spec.initContainers) == 0"
      message: "execution pods must not have init containers"   # the image and limit rules below check containers only
    - expression: "object.spec.containers.all(c, c.image.startsWith('ghcr.io/example/') && c.image.contains('@sha256:'))"
      message: "execution images must come from ghcr.io/example/ and be pinned by digest"
    - expression: "object.spec.containers.all(c, has(c.resources) && has(c.resources.limits) && 'cpu' in c.resources.limits && 'memory' in c.resources.limits)"
      message: "execution containers must set CPU and memory limits"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: code-exec-pod-shape
spec:
  policyName: code-exec-pod-shape
  validationActions: ["Deny"]
  matchResources:
    namespaceSelector:
      matchLabels:
        kubernetes.io/metadata.name: code-exec

4. In the executor: create the ConfigMap, then the pod; wait for it to finish or time out; read the log (cap its size); delete the pod and the ConfigMap; and on start-up, delete anything older than a few minutes left by a crash. Never reuse a pod, and never pass user input into fields of the pod spec other than the ConfigMap data.

Prove it

The policy's CEL was first evaluated with cel-python (standard CEL, outside the API server) against the page's pod and against pods that break one rule each. Real output:

text
ok   page pod: allow
ok   runc instead of gvisor-nonet: deny
ok   no runtimeClassName: deny
ok   token mounted: deny
ok   deadline 3600: deny
ok   restartPolicy Always: deny
ok   hostPath volume: deny
ok   image by tag: deny
ok   image from docker hub: deny
ok   no memory limit: deny

On a lab cluster (Kubernetes 1.34, gVisor release-20260921, with runsc-nonet configured), as the executor's ServiceAccount. For the run, the policy's registry prefix and the pod's image were pointed at a real image pinned by digest (docker.io/library/python@sha256:...), which is what you do with your own registry. Real output:

1. A privileged pod is refused:

bash
kubectl -n code-exec run evil --image=busybox:1.37 --privileged \
  --as=system:serviceaccount:executor:executor
text
Error from server (Forbidden): pods "evil" is forbidden: violates PodSecurity "restricted:v1.37": privileged (container "evil" must not set securityContext.privileged=true), allowPrivilegeEscalation != false ...

Pod Security stops it first; the policy would stop it next.

2. An image outside your registry, or not pinned by digest, is refused:

text
The pods "exec-9d41c2" is invalid: : ValidatingAdmissionPolicy 'code-exec-pod-shape' with binding 'code-exec-pod-shape' denied request: execution images must come from ghcr.io/example/ and be pinned by digest

3. One execution, end to end:

bash
kubectl -n code-exec create configmap exec-9d41c2 --from-literal=main.py="print(6 * 7)" --as=system:serviceaccount:executor:executor
kubectl apply -f k8s-exec-pod.yaml --as=system:serviceaccount:executor:executor
kubectl -n code-exec get pod exec-9d41c2 -o jsonpath='{.spec.runtimeClassName} {.spec.nodeName} {.status.phase}'
kubectl -n code-exec logs exec-9d41c2
text
configmap/exec-9d41c2 created
pod/exec-9d41c2 created
gvisor-nonet lab-worker Succeeded
42

Without the ConfigMap the pod never starts, and activeDeadlineSeconds: 30 ends it: DeadlineExceeded: Pod was active on the node longer than the specified deadline.

4. A run leaves nothing behind:

bash
kubectl -n code-exec delete pod,configmap exec-9d41c2 --as=system:serviceaccount:executor:executor
kubectl -n code-exec get pods,configmaps --no-headers | grep -v kube-root-ca.crt | wc -l
text
0

5. The executor cannot reach anything else:

bash
kubectl -n code-exec auth can-i list secrets --as=system:serviceaccount:executor:executor
kubectl -n default auth can-i create pods --as=system:serviceaccount:executor:executor
text
no
no

Mistakes people make

Warm pools

Reusing a pod or an interpreter across executions shares files, memory and environment between users. A fresh pod per execution costs pod start-up time; measure it on your nodes with the gVisor RuntimeClass, and pay it.

Trusting RBAC to shape pods

"Create pods" includes privileged pods, hostPath and any ServiceAccount in the namespace. Admission policy and Pod Security decide what those pods may look like.

Putting user input in the pod spec

A user-chosen name, image or environment value that reaches the pod spec is an injection point. Put user input in ConfigMap data only, and generate every other field.

Counting on garbage collection

Finished pods stay until the cluster crosses the terminated-pod threshold, which defaults to 12500 pods. Delete each pod and its ConfigMap when the run ends, and sweep on executor start-up.

Letting output grow without bound

A log is the output channel; a program can print gigabytes. Cap what the executor reads, and keep the ephemeral storage limit low.

Checklist

  • Every execution runs in a new pod that is deleted afterwards, with its ConfigMap.
  • Execution pods use a gVisor RuntimeClass with networking disabled.
  • The namespace enforces restricted Pod Security and holds no Secrets.
  • The executor's Role allows only pods, ConfigMaps and pod logs in that namespace.
  • A ValidatingAdmissionPolicy enforces RuntimeClass, no token, deadline, no restarts, allowed volumes, no init containers, pinned images and limits.
  • User input reaches the pod only as ConfigMap data.
  • The executor caps output size and sweeps leftovers on start-up.

Treat every execution like a guest who will not be coming back: a clean room, a locked door, a short visit, and fresh sheets for the next one.

H2-CSSE

Learn it on a live range

Secure design and threat modelling, in Secure Software: a real host in your browser, and every objective checked on the machine.

Start free

H2 Security services

Want it done with your team?

Our engineers set it up with you, test it the way this page does, and leave it documented.

See our services