Nodes and clusters

VPC segmentation for untrusted CI, the secure way

Your CI runner builds pull requests from strangers, runs postinstall scripts from ten thousand packages, and lives in the same private network as your etcd. What could a curl possibly do?

The short answer

Put CI runners in their own VPC with no peering to production, created with an explicit vpc_uuid so they never land in the default VPC. Give them no inbound rules, deny outbound traffic to private ranges and to your own Droplets by tag, allow only HTTPS and DNS out, and block job containers from the metadata service on the runner host.

Updated Houssam Hammoudi, CTOTested with DigitalOcean (VPC, Droplets, Cloud Firewalls API), doctl, nftables 1.0.2 and 1.1.3, Docker 29

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A CI runner is the one machine you invite attackers onto. It checks out pull requests, installs dependencies whose install scripts run with the job's rights, and executes test code nobody reviewed line by line. Every one of those is arbitrary code execution, by design.

What that code can reach is decided by the network the runner sits in. On most setups the answer is "everything": the runner was created in the region's default VPC, next to the cluster, the database and the bastion. From there a job can scan the private network, talk to etcd or the kubelet if a port is open, read the cloud metadata service, and reach any internal service that trusts "requests from inside".

Two defaults make this likely on DigitalOcean. A Droplet created without a vpc_uuid goes into the account's default VPC for the region, which is where production usually is. And even a runner in its own VPC is created with public networking unless you turn it off (public_networking defaults to on), so it can reach your cluster's public endpoints like any other host on the internet.

What the docs say

VPC networks are inaccessible from the public internet and other VPC networks, and traffic on them doesn't count against bandwidth usage.

Source: DigitalOcean docs, VPC

If excluded, the Droplet will be assigned to your account's default VPC for the region.

Source: DigitalOcean API reference, Droplets (vpc_uuid)

Deny rules take precedence over allow rules, including when the two rules belong to different firewalls applied to the same Droplet.

Source: DigitalOcean docs, How to Configure Firewall Rules

Firewalls affect both public and VPC network traffic.

Source: DigitalOcean docs, Cloud Firewalls limits

The VPC isolates private addresses only. The docs do not point out that a runner in a separate VPC still reaches your public endpoints, or that the metadata service at 169.254.169.254 answers any process on the runner, including job containers.

The secure configuration

The example uses region ams3, a CI VPC 10.20.0.0/20, the runner tag ci-runner, and production tags talos-cp, talos-worker and bastion.

1. A VPC for CI only, with no peering.

json
{
  "name": "ci-runners",
  "region": "ams3",
  "ip_range": "10.20.0.0/20",
  "description": "Untrusted CI jobs. No peering to any other VPC."
}

2. Runners created in that VPC, always with an explicit vpc_uuid.

json
{
  "name": "ci-runner-01",
  "region": "ams3",
  "size": "s-2vcpu-4gb",
  "image": "ubuntu-24-04-x64",
  "vpc_uuid": "3f2b6c1e-8d4a-4e2b-9c7f-1a2b3c4d5e6f",
  "tags": ["ci-runner"],
  "ipv6": false,
  "monitoring": false
}

Keep secrets out of user_data: jobs can read it from the metadata service. Register runners with a single-use token delivered at provisioning time, and rebuild runners instead of logging in to them.

3. A firewall with no inbound rules and denied egress to your networks. Deny rules win over allow rules, so the allow for HTTPS cannot reach your own Droplets or private ranges.

json
{
  "name": "ci-runner",
  "tags": ["ci-runner"],
  "outbound_rules": [
    {"protocol": "tcp", "ports": "0", "action": "deny",
     "destinations": {"addresses": ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "100.64.0.0/10"],
                      "tags": ["talos-cp", "talos-worker", "bastion"]}},
    {"protocol": "udp", "ports": "0", "action": "deny",
     "destinations": {"addresses": ["10.0.0.0/8", "172.16.0.0/12", "192.168.0.0/16", "100.64.0.0/10"],
                      "tags": ["talos-cp", "talos-worker", "bastion"]}},
    {"protocol": "tcp", "ports": "443", "action": "allow", "destinations": {"addresses": ["0.0.0.0/0", "::/0"]}},
    {"protocol": "udp", "ports": "53", "action": "allow", "destinations": {"addresses": ["1.1.1.1", "8.8.8.8"]}},
    {"protocol": "tcp", "ports": "53", "action": "allow", "destinations": {"addresses": ["1.1.1.1", "8.8.8.8"]}},
    {"protocol": "udp", "ports": "123", "action": "allow", "destinations": {"addresses": ["0.0.0.0/0"]}}
  ]
}

No inbound rules means no inbound traffic, not even SSH: the DigitalOcean docs say that if no inbound rules are configured, no incoming traffic is permitted. The firewall is stateful, so replies to the runner's own outbound connections still arrive. Runners pull jobs; nothing needs to connect to them. DNS is allowed only to 1.1.1.1 and 8.8.8.8, so point the runner's resolver (for example DNS= in systemd-resolved) at those addresses, or change both rules to the resolvers you use. The Droplet's default resolvers (the VPC resolver and DigitalOcean's) time out behind this firewall, and so does every docker pull until you change it. Port 80 is closed too, so package mirrors must be reached over HTTPS. Adding rules to a firewall does not end connections that are already open, so attach the firewall before the runner registers.

4. On the runner host, keep job containers away from the metadata service and from the host. The cloud firewall blocks traffic at the network layer before it reaches a Droplet. It never sees a container on the runner talking to the host itself, and the DigitalOcean docs do not say that it filters the link-local metadata address 169.254.169.254, so do not count on it; the first check under "Prove it" shows whether jobs can reach it.

conf
# ci-runner-host.nft
# Load with: nft -f /etc/nftables.d/ci-runner-host.nft (and include it from /etc/nftables.conf)
# One interface pattern per rule: older nft (1.0.2 in our test) rejects wildcards inside sets.
table inet ci_isolation {
  chain forward {
    type filter hook forward priority filter - 1; policy accept;
    # Traffic between job containers on the same bridge is not inspected here.
    iifname "docker0" oifname "docker0" accept
    iifname "br-*" oifname "br-*" accept
    # Traffic from job containers that leaves the bridge: no link-local
    # (metadata service) and no private or tailnet ranges.
    iifname "docker0" ip daddr { 169.254.0.0/16, 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10 } counter drop
    iifname "br-*" ip daddr { 169.254.0.0/16, 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 100.64.0.0/10 } counter drop
    iifname "docker0" ip6 daddr { fe80::/10, fc00::/7 } counter drop
    iifname "br-*" ip6 daddr { fe80::/10, fc00::/7 } counter drop
  }
  chain input {
    type filter hook input priority filter - 1; policy accept;
    # Job containers may not open connections to services on the host.
    iifname "docker0" ct state new counter drop
    iifname "br-*" ct state new counter drop
  }
}

The production side must hold too: nothing in production allows the internet, and therefore a runner, on 6443, 50000 or any admin port (see Cloud firewalls by tag).

Prove it

Built for real in ams3: the CI VPC, a runner Droplet in it (with Docker), and the firewall above, created through the API; DigitalOcean accepted the deny rules as written. A test Droplet outside the VPC stood in for production: its tag replaced the production tags in the deny rules. The firewall has no inbound rules, so a second, temporary firewall allowed SSH from one tester address to run the checks. The runner's resolver was set to 1.1.1.1 and 8.8.8.8.

1. The metadata service, from a job, before and after the host rules:

text
no host rules:        job -> metadata: 200   job -> host sshd: open     job -> https://proxy.golang.org/: 200
with the host rules:  job -> metadata: blocked   job -> host sshd: blocked   job -> https://proxy.golang.org/: 200
                      job on a user-defined bridge (br-*) -> metadata: blocked

The cloud firewall alone let jobs read the metadata service; the host rules stopped it. The host itself still reaches it (200), which the runner agent needs no more than the jobs do; keep secrets out of user_data.

The first version of the host rules on this page put the interface names in sets (iifname { "docker0", "br-*" }). nft 1.1.3 loaded it; the runner's own nft 1.0.2 refused it:

text
/root/ci-runner-host.nft:8:5-11: Error: Byteorder mismatch: expected big endian, got host endian

The file above uses one interface pattern per rule and loaded on both.

2. Production is out of reach, and the deny rule is what stops it:

text
                                          production :443   internet 1.1.1.1:443
deny rules with the production tag:       blocked           OPEN
control, the tag removed from deny rules: OPEN              OPEN

Port 443 is allowed to the internet, so only the tag in the deny rules keeps the runner away from production. Other checks from the runner, all blocked: another VPC's private address on 22, the production Droplet's public address on 22, the VPC gateway on 22, and 1.1.1.1:80.

3. DNS through the allowed resolvers only:

text
DNS via the VPC resolver    ;; communications error ... timed out
DNS via 67.207.67.2         ;; communications error ... timed out
DNS via 1.1.1.1             172.217.17.209

4. No peering includes the CI VPC:

bash
curl -sS -H "Authorization: Bearer $DO_READ_TOKEN" https://api.digitalocean.com/v2/vpc_peerings \
  | jq --arg vpc "$CI_VPC_ID" '[.vpc_peerings[] | select(.vpc_ids | index($vpc))] | length'
text
0

Run the same checks from a job on your runner, as a pipeline step. The request bodies above also validate against DigitalOcean's public OpenAPI spec (secure-tests/vpc-segmentation-untrusted-ci/).

Mistakes people make

Interface sets on an older nft

iifname { "docker0", "br-*" } fails to load on nft 1.0.2 with a byte order error, and a ruleset that does not load protects nothing. Write one rule per interface pattern, and check that the table exists after boot.

Creating runners without a vpc_uuid

The API puts the Droplet in the default VPC of the region, next to production. Make vpc_uuid required in whatever creates runners, and audit Droplets tagged ci-runner for the right VPC.

Treating a separate VPC as a separate world

VPC isolation covers private addresses. The runner can still reach every public address you have. Deny your own Droplets by tag in the runner firewall, and keep production's public ports closed to the internet.

Allowing all outbound "because builds need the internet"

Builds need HTTPS to registries and DNS. They do not need SSH, SMTP or random UDP. A narrow outbound list also slows down exfiltration and crypto mining.

Leaving the metadata service open to jobs

On DigitalOcean the metadata service serves the Droplet's user data. If a runner registration token or any other secret is there, every job can read it. Keep user data empty of secrets and block link-local addresses from job containers.

Long-lived runners

A runner that lives for months collects caches, credentials and whatever the last job left behind. Rebuild runners often, ideally one per job.

Checklist

  • CI runners live in their own VPC, and no VPC peering includes it.
  • Every runner is created with an explicit vpc_uuid.
  • The runner firewall has no inbound rules.
  • Outbound is denied to private ranges, the tailnet range and production tags, and allows only HTTPS, DNS and NTP.
  • Runner user data contains no secrets.
  • The runner host drops job container traffic to 169.254.0.0/16, private ranges and the host itself.
  • Production ports 6443, 50000 and admin ports are closed to the internet.
  • A pipeline step checks the metadata service, private targets and production ports on every runner image.
  • Runners are rebuilt regularly.

CI is remote code execution with a nice web interface. Give it a network where the worst thing it can reach is the package registry.

H2-CSDE

Learn it on a live range

Self-hosting the forge, in DevSecOps and Supply Chain: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on nodes and clusters

Talos Linux, Kubernetes API hardening, service account tokens, RBAC and the cloud underneath.

All nodes and clusters guides