One pod per code execution, the secure way
The "run this snippet" button was a hit, so the team kept a warm pool of interpreter pods to make it fast. The second user's code found the first user's variables still in memory, and a file called results.csv.
The short answer
Give every execution a new pod that is deleted afterwards: gVisor with no network, no ServiceAccount token, non-root, read-only root filesystem, input from a per-run ConfigMap, output through logs, a short deadline and hard limits. Let the executor create only such pods, enforced by a ValidatingAdmissionPolicy, in a namespace that holds nothing else.
On this page
What goes wrong
A feature that runs user code, a playground, an autograder, an AI agent's code tool, is remote code execution on purpose. The design decides how much of your system each execution touches.
The fast designs share things:
- Warm pools and long-lived interpreters. Execution B runs in the same process or pod as execution A, and finds A's files, environment and memory.
- The service's own pod. Code runs inside the API server that received it, next to its database credentials.
- A broad executor role. The service that creates execution pods can
create any pod in its namespace. If it is compromised, or tricked by input
it passes into a pod spec, it creates a privileged pod with a
hostPathmount and the node is gone. - Leftovers. Finished pods and their ConfigMaps stay around, with user code and output inside, until someone cleans up.
What the docs say
Permission to create workloads (either Pods, or workload resources that manage Pods) in a namespace implicitly grants access to many other resources in that namespace, such as Secrets, ConfigMaps, and PersistentVolumes that can be mounted in Pods.
Source: Kubernetes docs, Role Based Access Control Good Practices
Safely running untrusted code, such as when running third-party/user-provided code, or for software forensics.
Source: gVisor docs, Production guide (use cases)
For failed Pods, the API objects remain in the cluster's API until a human or controller process explicitly removes them.
Source: Kubernetes docs, Pod Lifecycle
RBAC can say "create pods", not "create pods that look like this", so the
shape of the pod has to be enforced by admission. And finished pods are only
garbage-collected once the cluster passes the terminated-pod-gc-threshold
of the controller manager, so cleanup is the executor's job.
The secure configuration
1. A namespace that holds executions and nothing else, restricted by Pod Security, labeled for the gVisor policy, with no Secrets.
# k8s-exec-namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
name: code-exec
labels:
sandbox.example.com/untrusted: "true"
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: v1.37
---
apiVersion: v1
kind: ServiceAccount
metadata:
name: default
namespace: code-exec
automountServiceAccountToken: false
---
# The executor runs elsewhere (namespace "executor") and may only manage pods,
# ConfigMaps and logs here.
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: executor
namespace: code-exec
rules:
- apiGroups: [""]
resources: ["pods", "configmaps"]
verbs: ["create", "get", "delete"]
- apiGroups: [""]
resources: ["pods/log"]
verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: executor
namespace: code-exec
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: executor
subjects:
- kind: ServiceAccount
name: executor
namespace: executor2. The pod the executor creates for each run. Code arrives in a ConfigMap made for this run; the result is the log.
# k8s-exec-pod.yaml
apiVersion: v1
kind: Pod
metadata:
name: exec-9d41c2
namespace: code-exec
labels:
app: code-exec
network.example.com/none: "true" # the clusterwide deny policy applies too
spec:
runtimeClassName: gvisor-nonet # gVisor with --network=none: only loopback
restartPolicy: Never
activeDeadlineSeconds: 30
terminationGracePeriodSeconds: 0
automountServiceAccountToken: false
enableServiceLinks: false
securityContext:
runAsNonRoot: true
runAsUser: 65534
runAsGroup: 65534
seccompProfile:
type: RuntimeDefault
containers:
- name: run
image: ghcr.io/example/python-runner@sha256:0000000000000000000000000000000000000000000000000000000000000000
command: ["python3", "-I", "/code/main.py"] # -I: ignore env vars and user site dirs
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
resources:
requests:
cpu: 250m
memory: 128Mi
ephemeral-storage: 64Mi
limits:
cpu: "1"
memory: 256Mi
ephemeral-storage: 128Mi
volumeMounts:
- name: code
mountPath: /code
readOnly: true
- name: tmp
mountPath: /tmp
volumes:
- name: code
configMap:
name: exec-9d41c2 # created for this run, deleted with the pod
- name: tmp
emptyDir:
sizeLimit: 64Mi3. Enforce that shape for everything the executor creates. RBAC lets the executor create pods; this policy decides which pods.
# k8s-exec-policy.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: code-exec-pod-shape
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE"]
resources: ["pods"]
validations:
- expression: "has(object.spec.runtimeClassName) && object.spec.runtimeClassName == 'gvisor-nonet'"
message: "execution pods must use the gvisor-nonet RuntimeClass"
- expression: "has(object.spec.automountServiceAccountToken) && object.spec.automountServiceAccountToken == false"
message: "execution pods must not mount a ServiceAccount token"
- expression: "has(object.spec.activeDeadlineSeconds) && object.spec.activeDeadlineSeconds <= 60"
message: "execution pods must set activeDeadlineSeconds of 60 or less"
- expression: "has(object.spec.restartPolicy) && object.spec.restartPolicy == 'Never'"
message: "execution pods must not restart"
- expression: "!has(object.spec.volumes) || object.spec.volumes.all(v, has(v.configMap) || has(v.emptyDir))"
message: "execution pods may mount only ConfigMaps and emptyDir"
- expression: "!has(object.spec.initContainers) || size(object.spec.initContainers) == 0"
message: "execution pods must not have init containers" # the image and limit rules below check containers only
- expression: "object.spec.containers.all(c, c.image.startsWith('ghcr.io/example/') && c.image.contains('@sha256:'))"
message: "execution images must come from ghcr.io/example/ and be pinned by digest"
- expression: "object.spec.containers.all(c, has(c.resources) && has(c.resources.limits) && 'cpu' in c.resources.limits && 'memory' in c.resources.limits)"
message: "execution containers must set CPU and memory limits"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: code-exec-pod-shape
spec:
policyName: code-exec-pod-shape
validationActions: ["Deny"]
matchResources:
namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: code-exec4. In the executor: create the ConfigMap, then the pod; wait for it to finish or time out; read the log (cap its size); delete the pod and the ConfigMap; and on start-up, delete anything older than a few minutes left by a crash. Never reuse a pod, and never pass user input into fields of the pod spec other than the ConfigMap data.
Prove it
The policy's CEL was first evaluated with cel-python (standard CEL, outside the API server) against the page's pod and against pods that break one rule each. Real output:
ok page pod: allow
ok runc instead of gvisor-nonet: deny
ok no runtimeClassName: deny
ok token mounted: deny
ok deadline 3600: deny
ok restartPolicy Always: deny
ok hostPath volume: deny
ok image by tag: deny
ok image from docker hub: deny
ok no memory limit: denyOn a lab cluster (Kubernetes 1.34, gVisor release-20260921, with
runsc-nonet configured), as the executor's ServiceAccount. For the run, the
policy's registry prefix and the pod's image were pointed at a real image
pinned by digest (docker.io/library/python@sha256:...), which is what you do
with your own registry. Real output:
1. A privileged pod is refused:
kubectl -n code-exec run evil --image=busybox:1.37 --privileged \
--as=system:serviceaccount:executor:executorError from server (Forbidden): pods "evil" is forbidden: violates PodSecurity "restricted:v1.37": privileged (container "evil" must not set securityContext.privileged=true), allowPrivilegeEscalation != false ...Pod Security stops it first; the policy would stop it next.
2. An image outside your registry, or not pinned by digest, is refused:
The pods "exec-9d41c2" is invalid: : ValidatingAdmissionPolicy 'code-exec-pod-shape' with binding 'code-exec-pod-shape' denied request: execution images must come from ghcr.io/example/ and be pinned by digest3. One execution, end to end:
kubectl -n code-exec create configmap exec-9d41c2 --from-literal=main.py="print(6 * 7)" --as=system:serviceaccount:executor:executor
kubectl apply -f k8s-exec-pod.yaml --as=system:serviceaccount:executor:executor
kubectl -n code-exec get pod exec-9d41c2 -o jsonpath='{.spec.runtimeClassName} {.spec.nodeName} {.status.phase}'
kubectl -n code-exec logs exec-9d41c2configmap/exec-9d41c2 created
pod/exec-9d41c2 created
gvisor-nonet lab-worker Succeeded
42Without the ConfigMap the pod never starts, and activeDeadlineSeconds: 30
ends it: DeadlineExceeded: Pod was active on the node longer than the specified deadline.
4. A run leaves nothing behind:
kubectl -n code-exec delete pod,configmap exec-9d41c2 --as=system:serviceaccount:executor:executor
kubectl -n code-exec get pods,configmaps --no-headers | grep -v kube-root-ca.crt | wc -l05. The executor cannot reach anything else:
kubectl -n code-exec auth can-i list secrets --as=system:serviceaccount:executor:executor
kubectl -n default auth can-i create pods --as=system:serviceaccount:executor:executorno
noMistakes people make
Warm pools
Reusing a pod or an interpreter across executions shares files, memory and environment between users. A fresh pod per execution costs pod start-up time; measure it on your nodes with the gVisor RuntimeClass, and pay it.
Trusting RBAC to shape pods
"Create pods" includes privileged pods, hostPath and any ServiceAccount in
the namespace. Admission policy and Pod Security decide what those pods may
look like.
Putting user input in the pod spec
A user-chosen name, image or environment value that reaches the pod spec is an injection point. Put user input in ConfigMap data only, and generate every other field.
Counting on garbage collection
Finished pods stay until the cluster crosses the terminated-pod threshold, which defaults to 12500 pods. Delete each pod and its ConfigMap when the run ends, and sweep on executor start-up.
Letting output grow without bound
A log is the output channel; a program can print gigabytes. Cap what the executor reads, and keep the ephemeral storage limit low.
Checklist
- Every execution runs in a new pod that is deleted afterwards, with its ConfigMap.
- Execution pods use a gVisor RuntimeClass with networking disabled.
- The namespace enforces
restrictedPod Security and holds no Secrets. - The executor's Role allows only pods, ConfigMaps and pod logs in that namespace.
- A ValidatingAdmissionPolicy enforces RuntimeClass, no token, deadline, no restarts, allowed volumes, no init containers, pinned images and limits.
- User input reaches the pod only as ConfigMap data.
- The executor caps output size and sweeps leftovers on start-up.
Treat every execution like a guest who will not be coming back: a clean room, a locked door, a short visit, and fresh sheets for the next one.
H2-CSSE
Learn it on a live range
Secure design and threat modelling, in Secure Software: a real host in your browser, and every objective checked on the machine.
Start freeH2 Security services
Want it done with your team?
Our engineers set it up with you, test it the way this page does, and leave it documented.
See our services