Sandboxing untrusted code

Running untrusted repositories in an isolated job, the secure way

The feature sounded harmless: paste a repository URL, get a test report. The first repository's Makefile read the job's token from disk, listed your Secrets and mailed them home before the tests even started.

The short answer

Run each untrusted repository in its own Kubernetes Job: gVisor runtime, no ServiceAccount token, non-root, read-only root filesystem, and an emptyDir for the checkout. Clone without submodules or hooks, allow egress only to the git host, and give the Job a deadline, no retries and a TTL so nothing outlives the run.

Updated Houssam Hammoudi, CTOTested with Kubernetes 1.34 (kind), gVisor release-20260921, Cilium 1.20.2, git 2.54.0 (alpine/git), curl 8.14.1

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Building, testing or even installing the dependencies of a repository runs code its author wrote: a Makefile, a setup.py, an npm postinstall script, a test that opens a socket. If the repository is untrusted, that is remote code execution, and the only question is what the code can reach.

A typical first version runs the work in the CI system's normal runner or in a long-lived worker pod:

  • The pod's ServiceAccount token is mounted, often with enough RBAC to create pods or read Secrets, because the worker "needs to manage jobs".
  • CI variables with registry and cloud credentials are in the environment.
  • The network is open: the internal API server, databases, the metadata service and the internet.
  • The worker lives on, so the next repository runs next to whatever the last one left behind.
  • There is no deadline, so a miner in a test suite runs until someone looks.

Cloning is not safe by default either. Submodules fetch more code from URLs the repository chooses, and some past Git vulnerabilities turned a recursive clone into code execution.

What the docs say

Safely running untrusted code, such as when running third-party/user-provided code, or for software forensics. Note: This guide is not appropriate for this use-case, and will instead focus on how to run an existing trusted stack with gVisor.

Source: gVisor docs, Production guide

Once a Job reaches activeDeadlineSeconds, all of its running Pods are terminated and the Job status will become type: Failed with reason: DeadlineExceeded.

Source: Kubernetes docs, Jobs

If a pod that is affected by a NetworkPolicy is created before the network plugin has completed NetworkPolicy handling, that pod may be started unprotected, and isolation rules will be applied when the NetworkPolicy handling is completed.

Source: Kubernetes docs, Network Policies

gVisor names running untrusted code as a use case, then says its production guide does not cover it. The Job docs cover deadlines and cleanup but assume the code in the Job is yours. The combination is left to you.

The secure configuration

1. A namespace just for these runs, labeled for the gVisor admission policy (see gVisor for untrusted workloads), with a ResourceQuota and no Secrets in it at all.

2. One network policy, created before any Job. Pods of a run may reach DNS and the git host over HTTPS, nothing else:

yaml
# cilium-untrusted-run.yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: untrusted-run
  namespace: untrusted-runs
spec:
  endpointSelector:
    matchLabels:
      app: untrusted-run
  ingress:
    - {}                                    # an empty rule: ingress default-deny, nothing connects to a run
  egress:
    - toEndpoints:
        - matchLabels:
            k8s:io.kubernetes.pod.namespace: kube-system
            k8s-app: kube-dns
      toPorts:
        - ports:
            - port: "53"
              protocol: ANY
          rules:
            dns:
              - matchName: git.example.com     # only this name resolves
    - toFQDNs:
        - matchName: git.example.com
      toPorts:
        - ports:
            - port: "443"
              protocol: TCP

An empty ingress: [] list is not enough: Cilium puts a pod into ingress default-deny only when a rule that selects it has an ingress section with at least one rule, and - {} is that rule. If builds need packages, add your internal package mirror by name, never the open internet.

3. One Job per run. The controller fills in the run ID and the URL:

yaml
# k8s-untrusted-run-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
  name: run-7f3a9c
  namespace: untrusted-runs
spec:
  backoffLimit: 0                     # a failure is a result, not a reason to run the code again
  activeDeadlineSeconds: 900          # 15 minutes, whatever the code does
  ttlSecondsAfterFinished: 300        # the Job and its pod are deleted 5 minutes after the end
  template:
    metadata:
      labels:
        app: untrusted-run
        run: "7f3a9c"
    spec:
      runtimeClassName: gvisor
      restartPolicy: Never
      automountServiceAccountToken: false
      enableServiceLinks: false
      dnsConfig:
        options:
          - name: ndots               # look up git.example.com as given, not via search domains
            value: "1"                # the DNS proxy refuses search-domain names; some images (musl) stop there
      securityContext:
        runAsNonRoot: true
        runAsUser: 10001
        runAsGroup: 10001
        fsGroup: 10001
        seccompProfile:
          type: RuntimeDefault
      initContainers:
        - name: fetch
          image: ghcr.io/example/git@sha256:0000000000000000000000000000000000000000000000000000000000000000
          args:
            - clone
            - --depth=1
            - --no-recurse-submodules       # no code from URLs the repository chooses
            - --config=core.hooksPath=/dev/null
            - --config=transfer.fsckObjects=true
            - https://git.example.com/org/repo.git
            - /work/src
          volumeMounts:
            - name: work
              mountPath: /work
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities:
              drop: ["ALL"]
      containers:
        - name: run
          image: ghcr.io/example/test-runner@sha256:0000000000000000000000000000000000000000000000000000000000000000
          workingDir: /work/src
          command: ["make", "test"]         # the result is the log; nothing else leaves the pod
          env:
            - name: HOME
              value: /work/home
          volumeMounts:
            - name: work
              mountPath: /work
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities:
              drop: ["ALL"]
          resources:
            requests:
              cpu: 500m
              memory: 512Mi
            limits:
              cpu: "2"
              memory: 2Gi
              ephemeral-storage: 4Gi
      volumes:
        - name: work
          emptyDir:
            sizeLimit: 4Gi

The controller that creates these Jobs holds the RBAC to do it; the Jobs hold none. It reads the result with kubectl logs job/run-7f3a9c and nothing else leaves the pod.

Prove it

Run on a lab cluster. The policy allowed github.com in place of git.example.com; the Job cloned https://github.com/octocat/Hello-World.git with alpine/git pinned by digest, and the run container was the curl image running a short script in place of make test. Both images ran as uid 10001 on the gvisor RuntimeClass.

1. The run: code fetched, no credentials, gVisor, only the git host.

text
$ kubectl -n untrusted-runs logs job/run-7f3a9c -c fetch
Cloning into '/work/src'...
$ kubectl -n untrusted-runs logs job/run-7f3a9c -c run
README
README: Hello World!
ls: /var/run/secrets/kubernetes.io/serviceaccount: No such file or directory
[   0.000000] Starting gVisor...
internet 000
git host 200

2. The flows:

text
$ hubble observe --namespace untrusted-runs --label run=7f3a9c --since 5m
untrusted-runs/run-7f3a9c-4qpjz:37757 (ID:9373) -> kube-system/coredns-66bc5c9577-snqbj:53 (ID:1734) dns-request proxy FORWARDED (DNS Query github.com. A)
untrusted-runs/run-7f3a9c-4qpjz:62136 (ID:9373) -> 140.82.121.4:443 (ID:16777219) policy-verdict:L3-L4 EGRESS ALLOWED (TCP Flags: SYN)
untrusted-runs/run-7f3a9c-4qpjz:25229 (ID:9373) -> kube-system/coredns-66bc5c9577-ssrvf:53 (ID:1734) dns-request proxy DROPPED (DNS Query example.com. A)

The attempt at example.com ended at the DNS proxy: no address, so no connection to drop.

3. Why ndots: 1. Two pods under the same policy, the curl image (musl), timing curl https://github.com:

text
Cilium reject code refused (the default):
  ndots 5 (the Kubernetes default):  000 exit 6 after 5s
  ndots 1:                           200 exit 0 after 1s
Cilium reject code nameError:
  ndots 5:                           200 exit 0 after 0s
  ndots 1:                           200 exit 0 after 0s

With the default ndots, the first lookup is github.com.untrusted-runs.svc.cluster.local. The DNS rule allows only github.com, so the proxy refuses it, and musl gives up on a refusal instead of trying the next name. ndots: 1 sends the real name first. dnsProxy.dnsRejectResponseCode: nameError also fixes it (see CI build pods); keep ndots: 1 anyway, so the Job does not depend on a cluster-wide setting.

4. The deadline and cleanup work. A copy of the Job running sleep 3600, with the timers shortened to activeDeadlineSeconds: 30 and ttlSecondsAfterFinished: 20:

text
t+33s: Failed reason=DeadlineExceeded
Job was active longer than specified deadline
t+52s: job gone
$ kubectl -n untrusted-runs get pod -l run=sleep
No resources found in untrusted-runs namespace.

5. The clone flags. In a container, against a local repository whose .gitmodules points at another repository:

text
cloned: src
submodule dir empty: yes
core.hooksPath=/dev/null
transfer.fsckobjects=true
control, --recurse-submodules: submodule fetched: yes

With the page's flags the submodule directory stays empty and hooks point at /dev/null; the control clone with --recurse-submodules fetched the submodule.

Mistakes people make

Reusing a worker pod

A long-lived worker runs repository B next to what repository A left in /tmp, in its caches and in its memory. One Job per run, deleted after.

Retrying failed runs

backoffLimit defaults to 6. For untrusted code, a retry is another round of the same attack. Set it to 0.

Mounting CI secrets "just for the private dependency"

Anything the job can read, the repository's code can read. Fetch private dependencies in a trusted step outside the Job, or not at all.

Opening the network to "the internet"

Build tools want the internet; attackers want it more. Allow the git host and your package mirror by name, and nothing else.

Creating the Job before the policy

A pod created before its network policy is handled can start with open networking. Create the policy once, when the namespace is created, and never delete it.

Checklist

  • Each run is its own Job in a namespace with no Secrets.
  • Pods run with runtimeClassName: gvisor, non-root, a read-only root filesystem and no ServiceAccount token.
  • The clone uses --no-recurse-submodules and core.hooksPath=/dev/null.
  • Egress allows DNS for the git host and HTTPS to the git host (and a package mirror), nothing else.
  • The network policy exists before the first Job.
  • Every Job sets backoffLimit: 0, activeDeadlineSeconds and ttlSecondsAfterFinished.
  • CPU, memory and ephemeral storage are limited.
  • Results leave through logs only.

You cannot review every repository someone hands you, but you can decide what its code wakes up next to. Make it an empty room with one window and a timer on the door.

H2-CSDE

Learn it on a live range

Policy gates and scanning, in DevSecOps and Supply Chain: a real host in your browser, and every objective checked on the machine.

Start free

H2 Security services

Want it done with your team?

Our engineers set it up with you, test it the way this page does, and leave it documented.

See our services