Sandboxing untrusted code

gVisor for untrusted workloads on Kubernetes, the secure way

You installed gVisor, created the RuntimeClass, and the demo pod printed "Starting gVisor". Then the real workload shipped without runtimeClassName, ran on plain runc with the host kernel, and nobody got an error, because nothing was wrong.

The short answer

Run gVisor on a dedicated, tainted node pool, and give its RuntimeClass a node selector and toleration. In namespaces for untrusted code, require runtimeClassName gvisor with a ValidatingAdmissionPolicy and forbid host namespaces and hostPath. Keep network policy and resource limits, and check each pod's dmesg for the gVisor boot log.

Updated Houssam Hammoudi, CTOTested with Kubernetes 1.34.0 (kind), gVisor release-20260921; Talos patch checked against the Talos 1.14 docs

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

gVisor runs a container on its own user-space kernel, the Sentry. The application's system calls go to the Sentry instead of the host kernel, so a kernel bug the application can trigger is, most of the time, a bug in the Sentry and not in the host. On Kubernetes you reach it through a RuntimeClass whose handler is runsc.

The protection is opt-in per pod, and every way of not opting in is quiet:

  • No runtimeClassName, no sandbox. The pod runs with the node's default runtime, runc. No warning, no event.
  • No scheduling on the RuntimeClass. Kubernetes then assumes every node supports the class. The scheduler can place the pod on a node without the handler, and the pod ends in the Failed phase there.
  • Host access still works. gVisor does not make hostNetwork, hostPID, hostPath or a privileged container safe; those hand the pod host resources directly.
  • The network is still the network. A sandboxed pod can reach every Service, the API server and the cloud metadata endpoint unless a network policy says otherwise.

On Talos, the gVisor system extension also needs unprivileged user namespaces. Talos sets user.max_user_namespaces to 0 by default, as the Kernel Self Protection Project (KSPP) recommends. Turning them on everywhere to run gVisor somewhere weakens every node.

What the docs say

If no runtimeClassName is specified, the default RuntimeHandler will be used, which is equivalent to the behavior when the RuntimeClass feature is disabled.

Source: Kubernetes docs, Runtime Class

A sandbox is not a substitute for a secure architecture.

Source: gVisor docs, Security Model

gVisor requires unprivileged user namespace creation, so Talos default setting should be overridden

Source: Sidero Labs extensions, gVisor README

runsc looks for gvisor-bin/ next to its own binary, so keep them together if you move them.

Source: gVisor docs, Installation

None of these pages tells you how to make the sandbox mandatory. Kubernetes documents the fallback to the default runtime as expected behavior, and the Talos extension README asks for the user namespace sysctl without saying to limit it to the nodes that run gVisor.

The secure configuration

1. Sandbox nodes. Build their image with the siderolabs/gvisor extension, then patch only those nodes:

yaml
# talos-gvisor-nodes.yaml
# Applied to the sandbox worker pool only.
apiVersion: v1alpha1
kind: SysctlConfig
params:
  user.max_user_namespaces: "11255"   # required by the gVisor extension; KSPP trade-off, so sandbox nodes only
---
apiVersion: v1alpha1
kind: KubeNodeConfig
labels:
  sandbox.example.com/gvisor: "true"
taints:
  sandbox.example.com/gvisor: NoSchedule   # "key: effect" (value empty); nothing else lands here by accident

Talos passes these taints to the kubelet, which adds them when the node first registers. The Talos reference warns that, with the default NodeRestriction admission plugin, worker nodes are not allowed to modify their taints after that. For a worker that has already joined, add the taint yourself: kubectl taint node <node> sandbox.example.com/gvisor:NoSchedule.

On other distributions, install runsc, containerd-shim-runsc-v1 and the gvisor-bin/ directory from the release tarball. runsc executes the sidecar binaries in gvisor-bin/ at runtime and looks for that directory next to its own binary. The fallback in runsc install that downloads missing files is scheduled to be dropped at the end of September 2026.

2. A RuntimeClass that only schedules where the handler exists.

yaml
# k8s-runtimeclass.yaml
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
  name: gvisor
handler: runsc
scheduling:
  nodeSelector:
    sandbox.example.com/gvisor: "true"   # merged into every pod that uses this class
  tolerations:
    - key: sandbox.example.com/gvisor
      operator: Exists
      effect: NoSchedule
overhead:
  podFixed:
    cpu: 100m        # room for the Sentry and gofer, added to the pod's resource
    memory: 64Mi     # accounting; example values, measure your own

3. Make the class mandatory where untrusted code runs.

yaml
# k8s-sandbox-policy.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: sandbox-requires-gvisor
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
      - apiGroups: [""]
        apiVersions: ["v1"]
        operations: ["CREATE", "UPDATE"]
        resources: ["pods"]
  validations:
    - expression: "has(object.spec.runtimeClassName) && object.spec.runtimeClassName in ['gvisor', 'gvisor-nonet']"
      message: "pods in sandbox namespaces must use a gVisor RuntimeClass (gvisor or gvisor-nonet)"   # list every gVisor class you run
    - expression: "!(has(object.spec.hostNetwork) && object.spec.hostNetwork) && !(has(object.spec.hostPID) && object.spec.hostPID) && !(has(object.spec.hostIPC) && object.spec.hostIPC)"
      message: "host namespaces are not allowed in sandbox namespaces"
    - expression: "!has(object.spec.volumes) || object.spec.volumes.all(v, !has(v.hostPath))"
      message: "hostPath volumes are not allowed in sandbox namespaces"
    - expression: "(has(object.spec.initContainers) ? object.spec.containers + object.spec.initContainers : object.spec.containers).all(c, !has(c.securityContext) || !has(c.securityContext.privileged) || !c.securityContext.privileged)"
      message: "privileged containers are not allowed in sandbox namespaces"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: sandbox-requires-gvisor
spec:
  policyName: sandbox-requires-gvisor
  validationActions: ["Deny"]
  matchResources:
    namespaceSelector:
      matchLabels:
        sandbox.example.com/untrusted: "true"
---
apiVersion: v1
kind: Namespace
metadata:
  name: sandbox
  labels:
    sandbox.example.com/untrusted: "true"
    pod-security.kubernetes.io/enforce: baseline      # never exempt the gvisor RuntimeClass from PSA
    pod-security.kubernetes.io/enforce-version: v1.37

4. The rest of the fence. Give the namespace a default-deny network policy (see Pods with no network at all), a ResourceQuota and LimitRange, and no ServiceAccount token (see ServiceAccount token hardening). gVisor's security model says network policy controls should be applied at the container level, outside the sandbox.

Prove it

Run on a lab cluster (Kubernetes 1.34, gVisor release-20260921) with the RuntimeClass and admission policy above; one worker was labeled and tainted the way the Talos patch does it. Real output.

1. The pod runs on gVisor's kernel:

bash
kubectl -n sandbox run gv --image=busybox:1.37 --restart=Never \
  --overrides='{"spec":{"runtimeClassName":"gvisor","automountServiceAccountToken":false}}' -- sleep 600
kubectl -n sandbox exec gv -- dmesg
text
[   0.000000] Starting gVisor...
[   0.405229] Rewriting operating system in Javascript...
[   0.424189] Singleplexing /dev/ptmx...
[   0.901604] Forking spaghetti code...
...
[   3.325443] Ready!

What it means: gVisor's boot log, starting with Starting gVisor... and followed by a few humorous lines, instead of the host kernel's messages (or a permission error). This is a quick check for the operator, not proof: the gVisor quick start notes that the output is easily replicated by an attacker, so code inside the pod must never use dmesg to decide whether it is sandboxed. Checks 3 and 4 below read the pod spec from the API instead.

2. The class is mandatory:

bash
kubectl -n sandbox run plain --image=busybox:1.37 --restart=Never -- sleep 600
text
The pods "plain" is invalid: : ValidatingAdmissionPolicy 'sandbox-requires-gvisor' with binding 'sandbox-requires-gvisor' denied request: pods in sandbox namespaces must use a gVisor RuntimeClass (gvisor or gvisor-nonet)

What it means: pods without a gVisor class never reach a node.

3. The pod landed on a sandbox node:

bash
kubectl -n sandbox get pod gv -o jsonpath='{.spec.runtimeClassName} {.spec.nodeName}{"\n"}'
kubectl get node -l sandbox.example.com/gvisor=true
text
gvisor lab-worker
NAME         STATUS   ROLES    AGE   VERSION
lab-worker   Ready    <none>   39m   v1.34.0

What it means: gvisor and a node name from the labeled pool.

4. Nothing in the sandbox namespace skips gVisor:

bash
kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}/{.metadata.name} {.spec.runtimeClassName}{"\n"}{end}' \
  | awk '$1 ~ /^sandbox\// && $2 != "gvisor"'
text
(no output)

What it means: no output.

Mistakes people make

One gVisor class in the policy, two in the cluster

A policy that demands runtimeClassName == 'gvisor' also rejects your other gVisor classes. On the lab, with this page's policy and the one pod per code execution guide both applied, every execution pod was refused, because that guide uses gvisor-nonet:

text
The pods "exec-9d41c2" is invalid: : ValidatingAdmissionPolicy 'sandbox-requires-gvisor' with binding 'sandbox-requires-gvisor' denied request: pods in sandbox namespaces must set runtimeClassName: gvisor

List every gVisor class you run in the expression, as the policy above does.

Relying on the pod spec alone

A RuntimeClass is a request, not a rule. Without an admission policy, one missing line puts untrusted code on runc. Require the class in every namespace where untrusted code runs.

Exempting the gVisor class from Pod Security

A PSA exemption for runtimeClassName: gvisor lets any pod that names the class skip every check, including hostPath and host namespaces. gVisor does not make those safe. Keep baseline enforced.

Enabling user namespaces on every node

The gVisor extension README asks for user.max_user_namespaces, and warns that it disables a KSPP setting. Apply it only to the sandbox pool, and keep other workloads off that pool with the taint.

Copying runsc without its helpers

Current gVisor releases ship runsc with a gvisor-bin/ directory of sidecar binaries that runsc executes at runtime. Install from the tarball, keep the directory next to runsc, and test a pod after every upgrade.

Believing the sandbox covers the network

gVisor has its own network stack, but packets still leave the pod. The metadata service, the API server and every Service are reachable until a network policy blocks them.

Checklist

  • gVisor runs on a labeled, tainted node pool, with its extension or full release tarball.
  • Sandbox nodes that joined before the patch carry the taint (kubectl describe node).
  • user.max_user_namespaces is raised on sandbox nodes only.
  • The gvisor RuntimeClass has scheduling.nodeSelector, tolerations and overhead.
  • A ValidatingAdmissionPolicy requires runtimeClassName: gvisor in untrusted namespaces.
  • The same policy forbids host namespaces, hostPath and privileged containers there.
  • Pod Security Admission exempts no RuntimeClass.
  • Untrusted namespaces have default-deny network policy, quotas and no ServiceAccount token.
  • kubectl exec ... dmesg in a sandbox pod shows the gVisor boot log.

gVisor is an excellent second kernel, but only for the pods that ask for it. Make asking mandatory, then check the kernel's diary.

H2-CSPE

Learn it on a live range

Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on sandboxing untrusted code

gVisor, pods with no network, and running other people's code without handing them your cluster.

All sandboxing untrusted code guides