gVisor for untrusted workloads on Kubernetes, the secure way
You installed gVisor, created the RuntimeClass, and the demo pod printed "Starting gVisor". Then the real workload shipped without runtimeClassName, ran on plain runc with the host kernel, and nobody got an error, because nothing was wrong.
The short answer
Run gVisor on a dedicated, tainted node pool, and give its RuntimeClass a node selector and toleration. In namespaces for untrusted code, require runtimeClassName gvisor with a ValidatingAdmissionPolicy and forbid host namespaces and hostPath. Keep network policy and resource limits, and check each pod's dmesg for the gVisor boot log.
On this page
What goes wrong
gVisor runs a container on its own user-space kernel, the Sentry. The
application's system calls go to the Sentry instead of the host kernel, so a
kernel bug the application can trigger is, most of the time, a bug in the
Sentry and not in the host. On Kubernetes you reach it through a
RuntimeClass whose handler is runsc.
The protection is opt-in per pod, and every way of not opting in is quiet:
- No
runtimeClassName, no sandbox. The pod runs with the node's default runtime, runc. No warning, no event. - No scheduling on the RuntimeClass. Kubernetes then assumes every node
supports the class. The scheduler can place the pod on a node without the
handler, and the pod ends in the
Failedphase there. - Host access still works. gVisor does not make
hostNetwork,hostPID,hostPathor a privileged container safe; those hand the pod host resources directly. - The network is still the network. A sandboxed pod can reach every Service, the API server and the cloud metadata endpoint unless a network policy says otherwise.
On Talos, the gVisor system extension also needs unprivileged user
namespaces. Talos sets user.max_user_namespaces to 0 by default, as the
Kernel Self Protection Project (KSPP) recommends. Turning them on everywhere to
run gVisor somewhere weakens every node.
What the docs say
If no runtimeClassName is specified, the default RuntimeHandler will be used, which is equivalent to the behavior when the RuntimeClass feature is disabled.
Source: Kubernetes docs, Runtime Class
A sandbox is not a substitute for a secure architecture.
Source: gVisor docs, Security Model
gVisor requires unprivileged user namespace creation, so Talos default setting should be overridden
Source: Sidero Labs extensions, gVisor README
runsc looks for gvisor-bin/ next to its own binary, so keep them together if you move them.
Source: gVisor docs, Installation
None of these pages tells you how to make the sandbox mandatory. Kubernetes documents the fallback to the default runtime as expected behavior, and the Talos extension README asks for the user namespace sysctl without saying to limit it to the nodes that run gVisor.
The secure configuration
1. Sandbox nodes. Build their image with the siderolabs/gvisor
extension, then patch only those nodes:
# talos-gvisor-nodes.yaml
# Applied to the sandbox worker pool only.
apiVersion: v1alpha1
kind: SysctlConfig
params:
user.max_user_namespaces: "11255" # required by the gVisor extension; KSPP trade-off, so sandbox nodes only
---
apiVersion: v1alpha1
kind: KubeNodeConfig
labels:
sandbox.example.com/gvisor: "true"
taints:
sandbox.example.com/gvisor: NoSchedule # "key: effect" (value empty); nothing else lands here by accidentTalos passes these taints to the kubelet, which adds them when the node first
registers. The Talos reference warns that, with the default NodeRestriction
admission plugin, worker nodes are not allowed to modify their taints after
that. For a worker that has already joined, add the taint yourself:
kubectl taint node <node> sandbox.example.com/gvisor:NoSchedule.
On other distributions, install runsc, containerd-shim-runsc-v1 and the
gvisor-bin/ directory from the release tarball. runsc executes the sidecar
binaries in gvisor-bin/ at runtime and looks for that directory next to its
own binary. The fallback in runsc install that downloads missing files is
scheduled to be dropped at the end of September 2026.
2. A RuntimeClass that only schedules where the handler exists.
# k8s-runtimeclass.yaml
apiVersion: node.k8s.io/v1
kind: RuntimeClass
metadata:
name: gvisor
handler: runsc
scheduling:
nodeSelector:
sandbox.example.com/gvisor: "true" # merged into every pod that uses this class
tolerations:
- key: sandbox.example.com/gvisor
operator: Exists
effect: NoSchedule
overhead:
podFixed:
cpu: 100m # room for the Sentry and gofer, added to the pod's resource
memory: 64Mi # accounting; example values, measure your own3. Make the class mandatory where untrusted code runs.
# k8s-sandbox-policy.yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: sandbox-requires-gvisor
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["pods"]
validations:
- expression: "has(object.spec.runtimeClassName) && object.spec.runtimeClassName in ['gvisor', 'gvisor-nonet']"
message: "pods in sandbox namespaces must use a gVisor RuntimeClass (gvisor or gvisor-nonet)" # list every gVisor class you run
- expression: "!(has(object.spec.hostNetwork) && object.spec.hostNetwork) && !(has(object.spec.hostPID) && object.spec.hostPID) && !(has(object.spec.hostIPC) && object.spec.hostIPC)"
message: "host namespaces are not allowed in sandbox namespaces"
- expression: "!has(object.spec.volumes) || object.spec.volumes.all(v, !has(v.hostPath))"
message: "hostPath volumes are not allowed in sandbox namespaces"
- expression: "(has(object.spec.initContainers) ? object.spec.containers + object.spec.initContainers : object.spec.containers).all(c, !has(c.securityContext) || !has(c.securityContext.privileged) || !c.securityContext.privileged)"
message: "privileged containers are not allowed in sandbox namespaces"
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: sandbox-requires-gvisor
spec:
policyName: sandbox-requires-gvisor
validationActions: ["Deny"]
matchResources:
namespaceSelector:
matchLabels:
sandbox.example.com/untrusted: "true"
---
apiVersion: v1
kind: Namespace
metadata:
name: sandbox
labels:
sandbox.example.com/untrusted: "true"
pod-security.kubernetes.io/enforce: baseline # never exempt the gvisor RuntimeClass from PSA
pod-security.kubernetes.io/enforce-version: v1.374. The rest of the fence. Give the namespace a default-deny network policy (see Pods with no network at all), a ResourceQuota and LimitRange, and no ServiceAccount token (see ServiceAccount token hardening). gVisor's security model says network policy controls should be applied at the container level, outside the sandbox.
Prove it
Run on a lab cluster (Kubernetes 1.34, gVisor release-20260921) with the RuntimeClass and admission policy above; one worker was labeled and tainted the way the Talos patch does it. Real output.
1. The pod runs on gVisor's kernel:
kubectl -n sandbox run gv --image=busybox:1.37 --restart=Never \
--overrides='{"spec":{"runtimeClassName":"gvisor","automountServiceAccountToken":false}}' -- sleep 600
kubectl -n sandbox exec gv -- dmesg[ 0.000000] Starting gVisor...
[ 0.405229] Rewriting operating system in Javascript...
[ 0.424189] Singleplexing /dev/ptmx...
[ 0.901604] Forking spaghetti code...
...
[ 3.325443] Ready!What it means: gVisor's boot log, starting with Starting gVisor...
and followed by a few humorous lines, instead of the host kernel's messages (or
a permission error). This is a quick check for the operator, not proof: the
gVisor quick start notes that the output is easily replicated by an attacker,
so code inside the pod must never use dmesg to decide whether it is
sandboxed. Checks 3 and 4 below read the pod spec from the API instead.
2. The class is mandatory:
kubectl -n sandbox run plain --image=busybox:1.37 --restart=Never -- sleep 600The pods "plain" is invalid: : ValidatingAdmissionPolicy 'sandbox-requires-gvisor' with binding 'sandbox-requires-gvisor' denied request: pods in sandbox namespaces must use a gVisor RuntimeClass (gvisor or gvisor-nonet)What it means: pods without a gVisor class never reach a node.
3. The pod landed on a sandbox node:
kubectl -n sandbox get pod gv -o jsonpath='{.spec.runtimeClassName} {.spec.nodeName}{"\n"}'
kubectl get node -l sandbox.example.com/gvisor=truegvisor lab-worker
NAME STATUS ROLES AGE VERSION
lab-worker Ready <none> 39m v1.34.0What it means: gvisor and a node name from the labeled pool.
4. Nothing in the sandbox namespace skips gVisor:
kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}/{.metadata.name} {.spec.runtimeClassName}{"\n"}{end}' \
| awk '$1 ~ /^sandbox\// && $2 != "gvisor"'(no output)What it means: no output.
Mistakes people make
One gVisor class in the policy, two in the cluster
A policy that demands runtimeClassName == 'gvisor' also rejects your other
gVisor classes. On the lab, with this page's policy and the
one pod per code execution guide
both applied, every execution pod was refused, because that guide uses
gvisor-nonet:
The pods "exec-9d41c2" is invalid: : ValidatingAdmissionPolicy 'sandbox-requires-gvisor' with binding 'sandbox-requires-gvisor' denied request: pods in sandbox namespaces must set runtimeClassName: gvisorList every gVisor class you run in the expression, as the policy above does.
Relying on the pod spec alone
A RuntimeClass is a request, not a rule. Without an admission policy, one missing line puts untrusted code on runc. Require the class in every namespace where untrusted code runs.
Exempting the gVisor class from Pod Security
A PSA exemption for runtimeClassName: gvisor lets any pod that names the
class skip every check, including hostPath and host namespaces. gVisor
does not make those safe. Keep baseline enforced.
Enabling user namespaces on every node
The gVisor extension README asks for user.max_user_namespaces, and warns
that it disables a KSPP setting. Apply it only to the sandbox pool, and keep
other workloads off that pool with the taint.
Copying runsc without its helpers
Current gVisor releases ship runsc with a gvisor-bin/ directory of
sidecar binaries that runsc executes at runtime. Install from the tarball,
keep the directory next to runsc, and test a pod after every upgrade.
Believing the sandbox covers the network
gVisor has its own network stack, but packets still leave the pod. The metadata service, the API server and every Service are reachable until a network policy blocks them.
Checklist
- gVisor runs on a labeled, tainted node pool, with its extension or full release tarball.
- Sandbox nodes that joined before the patch carry the taint (
kubectl describe node). user.max_user_namespacesis raised on sandbox nodes only.- The
gvisorRuntimeClass hasscheduling.nodeSelector,tolerationsandoverhead. - A ValidatingAdmissionPolicy requires
runtimeClassName: gvisorin untrusted namespaces. - The same policy forbids host namespaces,
hostPathand privileged containers there. - Pod Security Admission exempts no RuntimeClass.
- Untrusted namespaces have default-deny network policy, quotas and no ServiceAccount token.
kubectl exec ... dmesgin a sandbox pod shows the gVisor boot log.
gVisor is an excellent second kernel, but only for the pods that ask for it. Make asking mandatory, then check the kernel's diary.
H2-CSPE
Learn it on a live range
Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on sandboxing untrusted code
gVisor, pods with no network, and running other people's code without handing them your cluster.
All sandboxing untrusted code guides