Running untrusted repositories in an isolated job, the secure way
The feature sounded harmless: paste a repository URL, get a test report. The first repository's Makefile read the job's token from disk, listed your Secrets and mailed them home before the tests even started.
The short answer
Run each untrusted repository in its own Kubernetes Job: gVisor runtime, no ServiceAccount token, non-root, read-only root filesystem, and an emptyDir for the checkout. Clone without submodules or hooks, allow egress only to the git host, and give the Job a deadline, no retries and a TTL so nothing outlives the run.
On this page
What goes wrong
Building, testing or even installing the dependencies of a repository runs
code its author wrote: a Makefile, a setup.py, an npm postinstall
script, a test that opens a socket. If the repository is untrusted, that is
remote code execution, and the only question is what the code can reach.
A typical first version runs the work in the CI system's normal runner or in a long-lived worker pod:
- The pod's ServiceAccount token is mounted, often with enough RBAC to create pods or read Secrets, because the worker "needs to manage jobs".
- CI variables with registry and cloud credentials are in the environment.
- The network is open: the internal API server, databases, the metadata service and the internet.
- The worker lives on, so the next repository runs next to whatever the last one left behind.
- There is no deadline, so a miner in a test suite runs until someone looks.
Cloning is not safe by default either. Submodules fetch more code from URLs the repository chooses, and some past Git vulnerabilities turned a recursive clone into code execution.
What the docs say
Safely running untrusted code, such as when running third-party/user-provided code, or for software forensics. Note: This guide is not appropriate for this use-case, and will instead focus on how to run an existing trusted stack with gVisor.
Source: gVisor docs, Production guide
Once a Job reaches activeDeadlineSeconds, all of its running Pods are terminated and the Job status will become type: Failed with reason: DeadlineExceeded.
Source: Kubernetes docs, Jobs
If a pod that is affected by a NetworkPolicy is created before the network plugin has completed NetworkPolicy handling, that pod may be started unprotected, and isolation rules will be applied when the NetworkPolicy handling is completed.
Source: Kubernetes docs, Network Policies
gVisor names running untrusted code as a use case, then says its production guide does not cover it. The Job docs cover deadlines and cleanup but assume the code in the Job is yours. The combination is left to you.
The secure configuration
1. A namespace just for these runs, labeled for the gVisor admission policy (see gVisor for untrusted workloads), with a ResourceQuota and no Secrets in it at all.
2. One network policy, created before any Job. Pods of a run may reach DNS and the git host over HTTPS, nothing else:
# cilium-untrusted-run.yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: untrusted-run
namespace: untrusted-runs
spec:
endpointSelector:
matchLabels:
app: untrusted-run
ingress:
- {} # an empty rule: ingress default-deny, nothing connects to a run
egress:
- toEndpoints:
- matchLabels:
k8s:io.kubernetes.pod.namespace: kube-system
k8s-app: kube-dns
toPorts:
- ports:
- port: "53"
protocol: ANY
rules:
dns:
- matchName: git.example.com # only this name resolves
- toFQDNs:
- matchName: git.example.com
toPorts:
- ports:
- port: "443"
protocol: TCPAn empty ingress: [] list is not enough: Cilium puts a pod into ingress
default-deny only when a rule that selects it has an ingress section with at
least one rule, and - {} is that rule. If builds need packages, add your
internal package mirror by name, never the open internet.
3. One Job per run. The controller fills in the run ID and the URL:
# k8s-untrusted-run-job.yaml
apiVersion: batch/v1
kind: Job
metadata:
name: run-7f3a9c
namespace: untrusted-runs
spec:
backoffLimit: 0 # a failure is a result, not a reason to run the code again
activeDeadlineSeconds: 900 # 15 minutes, whatever the code does
ttlSecondsAfterFinished: 300 # the Job and its pod are deleted 5 minutes after the end
template:
metadata:
labels:
app: untrusted-run
run: "7f3a9c"
spec:
runtimeClassName: gvisor
restartPolicy: Never
automountServiceAccountToken: false
enableServiceLinks: false
dnsConfig:
options:
- name: ndots # look up git.example.com as given, not via search domains
value: "1" # the DNS proxy refuses search-domain names; some images (musl) stop there
securityContext:
runAsNonRoot: true
runAsUser: 10001
runAsGroup: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
initContainers:
- name: fetch
image: ghcr.io/example/git@sha256:0000000000000000000000000000000000000000000000000000000000000000
args:
- clone
- --depth=1
- --no-recurse-submodules # no code from URLs the repository chooses
- --config=core.hooksPath=/dev/null
- --config=transfer.fsckObjects=true
- https://git.example.com/org/repo.git
- /work/src
volumeMounts:
- name: work
mountPath: /work
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
containers:
- name: run
image: ghcr.io/example/test-runner@sha256:0000000000000000000000000000000000000000000000000000000000000000
workingDir: /work/src
command: ["make", "test"] # the result is the log; nothing else leaves the pod
env:
- name: HOME
value: /work/home
volumeMounts:
- name: work
mountPath: /work
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: "2"
memory: 2Gi
ephemeral-storage: 4Gi
volumes:
- name: work
emptyDir:
sizeLimit: 4GiThe controller that creates these Jobs holds the RBAC to do it; the Jobs hold
none. It reads the result with kubectl logs job/run-7f3a9c and nothing else
leaves the pod.
Prove it
Run on a lab cluster. The policy allowed github.com in place of
git.example.com; the Job cloned https://github.com/octocat/Hello-World.git
with alpine/git pinned by digest, and the run container was the curl image
running a short script in place of make test. Both images ran as uid 10001
on the gvisor RuntimeClass.
1. The run: code fetched, no credentials, gVisor, only the git host.
$ kubectl -n untrusted-runs logs job/run-7f3a9c -c fetch
Cloning into '/work/src'...
$ kubectl -n untrusted-runs logs job/run-7f3a9c -c run
README
README: Hello World!
ls: /var/run/secrets/kubernetes.io/serviceaccount: No such file or directory
[ 0.000000] Starting gVisor...
internet 000
git host 2002. The flows:
$ hubble observe --namespace untrusted-runs --label run=7f3a9c --since 5m
untrusted-runs/run-7f3a9c-4qpjz:37757 (ID:9373) -> kube-system/coredns-66bc5c9577-snqbj:53 (ID:1734) dns-request proxy FORWARDED (DNS Query github.com. A)
untrusted-runs/run-7f3a9c-4qpjz:62136 (ID:9373) -> 140.82.121.4:443 (ID:16777219) policy-verdict:L3-L4 EGRESS ALLOWED (TCP Flags: SYN)
untrusted-runs/run-7f3a9c-4qpjz:25229 (ID:9373) -> kube-system/coredns-66bc5c9577-ssrvf:53 (ID:1734) dns-request proxy DROPPED (DNS Query example.com. A)The attempt at example.com ended at the DNS proxy: no address, so no
connection to drop.
3. Why ndots: 1. Two pods under the same policy, the curl image (musl),
timing curl https://github.com:
Cilium reject code refused (the default):
ndots 5 (the Kubernetes default): 000 exit 6 after 5s
ndots 1: 200 exit 0 after 1s
Cilium reject code nameError:
ndots 5: 200 exit 0 after 0s
ndots 1: 200 exit 0 after 0sWith the default ndots, the first lookup is
github.com.untrusted-runs.svc.cluster.local. The DNS rule allows only
github.com, so the proxy refuses it, and musl gives up on a refusal instead
of trying the next name. ndots: 1 sends the real name first.
dnsProxy.dnsRejectResponseCode: nameError also fixes it (see
CI build pods); keep ndots: 1
anyway, so the Job does not depend on a cluster-wide setting.
4. The deadline and cleanup work. A copy of the Job running
sleep 3600, with the timers shortened to activeDeadlineSeconds: 30 and
ttlSecondsAfterFinished: 20:
t+33s: Failed reason=DeadlineExceeded
Job was active longer than specified deadline
t+52s: job gone
$ kubectl -n untrusted-runs get pod -l run=sleep
No resources found in untrusted-runs namespace.5. The clone flags. In a container, against a local repository whose
.gitmodules points at another repository:
cloned: src
submodule dir empty: yes
core.hooksPath=/dev/null
transfer.fsckobjects=true
control, --recurse-submodules: submodule fetched: yesWith the page's flags the submodule directory stays empty and hooks point at
/dev/null; the control clone with --recurse-submodules fetched the
submodule.
Mistakes people make
Reusing a worker pod
A long-lived worker runs repository B next to what repository A left in
/tmp, in its caches and in its memory. One Job per run, deleted after.
Retrying failed runs
backoffLimit defaults to 6. For untrusted code, a retry is another round of
the same attack. Set it to 0.
Mounting CI secrets "just for the private dependency"
Anything the job can read, the repository's code can read. Fetch private dependencies in a trusted step outside the Job, or not at all.
Opening the network to "the internet"
Build tools want the internet; attackers want it more. Allow the git host and your package mirror by name, and nothing else.
Creating the Job before the policy
A pod created before its network policy is handled can start with open networking. Create the policy once, when the namespace is created, and never delete it.
Checklist
- Each run is its own Job in a namespace with no Secrets.
- Pods run with
runtimeClassName: gvisor, non-root, a read-only root filesystem and no ServiceAccount token. - The clone uses
--no-recurse-submodulesandcore.hooksPath=/dev/null. - Egress allows DNS for the git host and HTTPS to the git host (and a package mirror), nothing else.
- The network policy exists before the first Job.
- Every Job sets
backoffLimit: 0,activeDeadlineSecondsandttlSecondsAfterFinished. - CPU, memory and ephemeral storage are limited.
- Results leave through logs only.
You cannot review every repository someone hands you, but you can decide what its code wakes up next to. Make it an empty room with one window and a timer on the door.
H2-CSDE
Learn it on a live range
Policy gates and scanning, in DevSecOps and Supply Chain: a real host in your browser, and every objective checked on the machine.
Start freeH2 Security services
Want it done with your team?
Our engineers set it up with you, test it the way this page does, and leave it documented.
See our services