Identity and access

A userspace Tailscale subnet router on Kubernetes, the secure way

The subnet router guide asks for NET_ADMIN, the auto-approver says 10.0.0.0/8, and the policy says the whole Service CIDR. Congratulations, you have built a VPN into every service in the cluster.

The short answer

Run the router in userspace mode as a non-root pod with all capabilities dropped and a read-only root. Tag it, auto-approve only its exact Service CIDR for that tag, and grant each group access to named services and ports, never the whole range. Give it a Role for its own state Secret only.

Updated Houssam Hammoudi, CTOTested with Headscale 0.29.4, Tailscale 1.102.4, Cilium 1.20.2, Kubernetes 1.34 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

The quick way to reach cluster services from laptops is a subnet router pod that advertises the Service CIDR. In kernel mode it needs /dev/net/tun and extra capabilities such as NET_ADMIN. The subnet router example in Tailscale's repository goes further: TS_USERSPACE=false and privileged: true. A compromised router pod then owns part of the node's network stack.

The route gets approved by hand once, or by an auto-approver for a wide range such as 10.0.0.0/8 owned by a group of people. Then any of their laptops that advertises the range is approved too, and can pull traffic meant for the cluster.

The grant is "group ops may reach 10.96.0.0/16". That includes the API server and every internal service, most of which never expected a laptop as a client.

The router stores its node key in a Secret. The example Role in Tailscale's repository also grants create on Secrets, and create cannot be limited to one name, so the pod may create any Secret in the namespace.

What the docs say

Userspace networking reduces the permissions Tailscale needs to run.

Source: Tailscale docs, Kubernetes

By default, without a policy enabled, all nodes can accept and use such routes.

Source: Headscale docs, Routes

By default, when you advertise subnet routes, Tailscale uses source network address translation (SNAT) (also called masquerading).

Source: Tailscale docs, Subnet routers

SNAT means every tailnet user reaches the cluster from the router pod's address. Kubernetes NetworkPolicy cannot tell them apart. Access control for people has to live in the tailnet policy; the cluster can only limit the router as a whole.

The secure configuration

In the Headscale (or Tailscale) policy: a tag for the router, an exact auto-approved route for that tag, and grants per service:

json
{
  "groups": {
    "group:ops": ["[email protected]"],
    "group:dev": ["[email protected]"],
  },
  "tagOwners": {
    "tag:k8s-router": ["group:ops"],
  },
  "hosts": {
    "cluster-services": "10.96.0.0/16",       // the cluster's Service CIDR (example)
    "grafana":          "10.96.20.5/32",
    "argocd":           "10.96.30.7/32",
    "kube-api":         "10.96.0.1/32",
  },
  // Only a node tagged tag:k8s-router gets this route (or a subnet of it) approved automatically.
  "autoApprovers": {
    "routes": {
      "10.96.0.0/16": ["tag:k8s-router"],
    },
  },
  "grants": [
    {"src": ["group:ops"], "dst": ["grafana", "argocd"], "ip": ["tcp:443"]},
    {"src": ["group:dev"], "dst": ["grafana"], "ip": ["tcp:443"]},
    // nobody gets the router itself, and nobody gets the whole Service CIDR
  ],
  "tests": [
    {"src": "[email protected]", "accept": ["grafana:443", "argocd:443"], "deny": ["kube-api:443"]},
    {"src": "[email protected]",   "accept": ["grafana:443"], "deny": ["argocd:443", "grafana:22"]},
  ],
}

In the cluster: a ServiceAccount whose Role reaches only its own state Secret, a non-root userspace router, and a NetworkPolicy for the router pod:

yaml
apiVersion: v1
kind: ServiceAccount
metadata:
  name: ts-router
  namespace: tailscale
---
# The state Secret is created empty up front, so the Role never needs "create".
apiVersion: v1
kind: Secret
metadata:
  name: ts-router-state
  namespace: tailscale
type: Opaque
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: ts-router-state
  namespace: tailscale
rules:
  - apiGroups: [""]
    resources: ["secrets"]
    resourceNames: ["ts-router-state"]          # its own state, nothing else
    verbs: ["get", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: ts-router-state
  namespace: tailscale
subjects:
  - kind: ServiceAccount
    name: ts-router
    namespace: tailscale
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: Role
  name: ts-router-state
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ts-router
  namespace: tailscale
spec:
  replicas: 1
  strategy:
    type: Recreate                                # one node identity per state Secret
  selector:
    matchLabels: {app: ts-router}
  template:
    metadata:
      labels: {app: ts-router}
    spec:
      serviceAccountName: ts-router
      securityContext:
        runAsNonRoot: true
        runAsUser: 1000
        runAsGroup: 1000
        seccompProfile: {type: RuntimeDefault}
      containers:
        - name: tailscale
          image: tailscale/tailscale:v1.102.4       # pin; better, by digest
          env:
            - {name: TS_USERSPACE, value: "true"}   # no TUN device, no NET_ADMIN
            - {name: TS_KUBE_SECRET, value: ts-router-state}
            - {name: TS_ROUTES, value: "10.96.0.0/16"}
            - {name: TS_EXTRA_ARGS, value: "--login-server=https://hs.example.com"}   # the tagged auth key sets the tag
            - name: TS_AUTHKEY                      # single-use, tagged, minutes; synced from OpenBao
              valueFrom:
                secretKeyRef: {name: ts-router-authkey, key: authkey, optional: true}
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities:
              drop: [ALL]
          resources:
            requests: {cpu: 50m, memory: 64Mi}
            limits: {memory: 256Mi}
          volumeMounts:
            - {name: tmp, mountPath: /tmp}
      volumes:
        - name: tmp
          emptyDir: {}
---
# The router may reach the two services it fronts, DNS, and the tailnet
# (control server, relays, peers). Nothing else in the cluster.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: ts-router
  namespace: tailscale
spec:
  podSelector:
    matchLabels: {app: ts-router}
  policyTypes: [Ingress, Egress]
  ingress: []
  egress:
    - to:                                           # the API server, for its state Secret
        - ipBlock: {cidr: 192.0.2.10/32}            # your API server endpoint(s): kubectl get endpoints kubernetes
      ports:
        - {protocol: TCP, port: 6443}
    - to:
        - namespaceSelector:
            matchLabels: {kubernetes.io/metadata.name: kube-system}
          podSelector:
            matchLabels: {k8s-app: kube-dns}
      ports:
        - {protocol: UDP, port: 53}
        - {protocol: TCP, port: 53}
    - to:
        - namespaceSelector:
            matchLabels: {kubernetes.io/metadata.name: monitoring}
          podSelector:
            matchLabels: {app.kubernetes.io/name: grafana}
      ports:
        - {protocol: TCP, port: 3000}
    - to:
        - namespaceSelector:
            matchLabels: {kubernetes.io/metadata.name: argocd}
          podSelector:
            matchLabels: {app.kubernetes.io/name: argocd-server}
      ports:
        - {protocol: TCP, port: 8080}
    - to:
        - ipBlock:
            cidr: 0.0.0.0/0
            except: [10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16]
      ports:
        - {protocol: TCP, port: 443}                # control server and DERP
        - {protocol: UDP, port: 3478}               # STUN
        - {protocol: UDP, port: 41641}              # peers on the default port; others relay via DERP

The router reads and writes its state Secret through the Kubernetes API, so it needs egress to the API server; without that rule it never starts tailscaled. On Cilium, the API server has its own identity and an ipBlock for its address did not let the pod through in our test; use a CiliumNetworkPolicy next to the NetworkPolicy:

yaml
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: ts-router-apiserver
  namespace: tailscale
spec:
  endpointSelector: {matchLabels: {app: ts-router}}
  egress:
    - toEntities: [kube-apiserver]
      toPorts: [{ports: [{port: "6443", protocol: TCP}]}]

With Headscale 0.29.4, do not add --advertise-tags=tag:k8s-router to a node that joins with a tagged auth key: Headscale refuses the registration with requested tags [tag:k8s-router] are invalid or not permitted. The key already makes the node tag:k8s-router.

After the router has joined, delete the ts-router-authkey Secret. The node key lives in ts-router-state from then on. With these variables, the image's start program (containerboot, v1.102.4 source) runs tailscaled with --tun=userspace-networking, --state=kube:ts-router-state and --statedir=/tmp, and puts its socket at /tmp/tailscaled.sock. That is why /tmp is the only writable mount. The Docker parameters page gives /var/run/tailscale/tailscaled.sock as the socket default; if a later image follows the page, set TS_SOCKET=/tmp/tailscaled.sock. Tailscale's Docker parameters page recommends TS_AUTH_ONCE=true when state is persisted, so a restart does not try to log in again.

Tailscale's policy file reference says an auto-approver for a route may also advertise a subnet of it. Here that means tag:k8s-router could get 10.96.20.0/24 approved, never anything outside 10.96.0.0/16. Headscale's routes page does not say either way; check with headscale nodes list-routes after a change.

Prove it

From secure-tests/userspace-tailscale-subnet-router-kubernetes/. Headscale 0.29.4 runs in a container. Tailscale clients join with the same restrictions as the pod above: uid 1000, --cap-drop ALL, a read-only root and userspace networking. The router advertises the Service CIDR and, by mistake, 10.0.0.0/8. A laptop advertises the Service CIDR too:

text
k8s-router: joined as uid 1000, no capabilities
Some peers are advertising routes but --accept-routes is false
bob-laptop: joined as uid 1000, no capabilities

Only the exact route, from the tagged router, is approved:

bash
headscale nodes list-routes
text
k8s-router  approved=10.96.0.0/16  available=10.0.0.0/8,10.96.0.0/16
bob-laptop  approved=              available=10.96.0.0/16

The grants and their tests pass: ops reaches Grafana and Argo CD on 443, dev reaches Grafana only, and nobody reaches the API server through the router:

bash
headscale policy check -f policy.hujson
text
Policy is valid

The manifests pass schema validation:

text
Summary: 6 resources found in 1 file - Valid: 6, Invalid: 0, Errors: 0, Skipped: 0

In a lab cluster, with the Deployment, RBAC and policies above. Headscale ran next to the cluster over HTTPS with a lab certificate (mounted into the router and the laptops through SSL_CERT_FILE), and the NetworkPolicy got two lab rules for Headscale and the laptops, which sat on a private network standing in for the internet. Stand-ins answered at the Grafana (10.96.20.5) and Argo CD (10.96.30.7) Service addresses. Two userspace Tailscale clients joined as [email protected] (ops) and [email protected] (dev), with --accept-routes.

What each person reaches through the router:

text
$ headscale nodes list-routes
ts-router-86595576ff-zlnxh  tags=tag:k8s-router  approved=10.96.0.0/16  available=10.96.0.0/16
alice (ops)  grafana: 200  argocd: 200      kube-api: timeout
bob   (dev)  grafana: 200  argocd: timeout  kube-api: timeout

What went wrong on the way:

text
NetworkPolicy without the API server rule:
  boot: error setting up for running on Kubernetes: getting Tailscale state Secret ts-router-state:
        Get "https://kubernetes.default.svc/api/v1/namespaces/tailscale/secrets/ts-router-state": context deadline exceeded
tagged auth key plus --advertise-tags=tag:k8s-router:
  Received error: handling register with auth key: creating new node: requested tags [tag:k8s-router] are invalid or not permitted

The router pod:

text
$ kubectl -n tailscale exec deploy/ts-router -- id
uid=1000 gid=1000 groups=1000
$ kubectl auth can-i list secrets -n tailscale --as=system:serviceaccount:tailscale:ts-router
no
$ kubectl auth can-i get secret/ts-router-state -n tailscale --as=system:serviceaccount:tailscale:ts-router
yes
$ kubectl auth can-i get secret/ts-router-authkey -n tailscale --as=system:serviceaccount:tailscale:ts-router
no
router -> grafana (in the policy):                     200
router -> cert-manager webhook (not in the policy):    timed out

After ts-router-authkey was deleted and the pod restarted, the router came back as the same node from ts-router-state.

Mistakes people make

No egress to the API server

The router keeps its identity in a Secret and needs the API server for it. A NetworkPolicy without that rule leaves the pod running and the tailnet empty.

Privileged pods for a router that does not need them

Userspace mode needs no TUN device and no capabilities, and it is the image's default (TS_USERSPACE defaults to true). The test ran the client as uid 1000 with every capability dropped.

Auto-approving a range for anyone

An auto-approver for 10.0.0.0/8 owned by a group lets any member's laptop become a router for your whole private space. Approve exact CIDRs for a tag that only ops can assign.

Granting the whole Service CIDR

A grant to 10.96.0.0/16 includes the API server, DNS and every internal service. Name the services and ports.

Relying on NetworkPolicy to tell users apart

With SNAT, all tailnet traffic comes from the router pod. The cluster can restrict the router; only the tailnet policy can restrict people.

A Role with create and list on Secrets

The state Secret can be created ahead of time. Then get, update and patch with resourceNames are enough.

A reusable auth key in a Secret forever

Use a single-use, tagged key that expires in minutes, and delete it once the router has joined.

Checklist

  • Set TS_USERSPACE=true; drop all capabilities; run as non-root with a read-only root.
  • Tag the router and restrict who may assign the tag.
  • Auto-approve only the Service CIDR (subnets of it are approved too), only for the router's tag.
  • Grant groups access to named service IPs and ports, never the whole CIDR.
  • Add policy tests that deny the API server and other internal services.
  • Pre-create the state Secret; give the Role get, update, patch on it by name.
  • Use a single-use, short-lived, tagged auth key and delete it after the first join.
  • Limit the router's egress with a NetworkPolicy.

A subnet router is a door into the cluster. Make it a narrow one with a list of rooms, not a hole in the wall.

H2-CPQE

Learn it on a live range

Private access and tailnets, in Edge and Post-Quantum Networking: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on identity and access

Self-hosted identity with Zitadel, private access with Headscale, break-glass and offboarding.

All identity and access guides