A userspace Tailscale subnet router on Kubernetes, the secure way
The subnet router guide asks for NET_ADMIN, the auto-approver says 10.0.0.0/8, and the policy says the whole Service CIDR. Congratulations, you have built a VPN into every service in the cluster.
The short answer
Run the router in userspace mode as a non-root pod with all capabilities dropped and a read-only root. Tag it, auto-approve only its exact Service CIDR for that tag, and grant each group access to named services and ports, never the whole range. Give it a Role for its own state Secret only.
On this page
What goes wrong
The quick way to reach cluster services from laptops is a subnet router pod
that advertises the Service CIDR. In kernel mode it needs /dev/net/tun and
extra capabilities such as NET_ADMIN. The subnet router example in
Tailscale's repository goes further: TS_USERSPACE=false and
privileged: true.
A compromised router pod then owns part of the node's network stack.
The route gets approved by hand once, or by an auto-approver for a wide range
such as 10.0.0.0/8 owned by a group of people. Then any of their laptops that
advertises the range is approved too, and can pull traffic meant for the
cluster.
The grant is "group ops may reach 10.96.0.0/16". That includes the API server and every internal service, most of which never expected a laptop as a client.
The router stores its node key in a Secret. The example Role in Tailscale's
repository also grants create on Secrets, and create cannot be limited to
one name, so the pod may create any Secret in the namespace.
What the docs say
Userspace networking reduces the permissions Tailscale needs to run.
Source: Tailscale docs, Kubernetes
By default, without a policy enabled, all nodes can accept and use such routes.
Source: Headscale docs, Routes
By default, when you advertise subnet routes, Tailscale uses source network address translation (SNAT) (also called masquerading).
Source: Tailscale docs, Subnet routers
SNAT means every tailnet user reaches the cluster from the router pod's address. Kubernetes NetworkPolicy cannot tell them apart. Access control for people has to live in the tailnet policy; the cluster can only limit the router as a whole.
The secure configuration
In the Headscale (or Tailscale) policy: a tag for the router, an exact auto-approved route for that tag, and grants per service:
{
"groups": {
"group:ops": ["[email protected]"],
"group:dev": ["[email protected]"],
},
"tagOwners": {
"tag:k8s-router": ["group:ops"],
},
"hosts": {
"cluster-services": "10.96.0.0/16", // the cluster's Service CIDR (example)
"grafana": "10.96.20.5/32",
"argocd": "10.96.30.7/32",
"kube-api": "10.96.0.1/32",
},
// Only a node tagged tag:k8s-router gets this route (or a subnet of it) approved automatically.
"autoApprovers": {
"routes": {
"10.96.0.0/16": ["tag:k8s-router"],
},
},
"grants": [
{"src": ["group:ops"], "dst": ["grafana", "argocd"], "ip": ["tcp:443"]},
{"src": ["group:dev"], "dst": ["grafana"], "ip": ["tcp:443"]},
// nobody gets the router itself, and nobody gets the whole Service CIDR
],
"tests": [
{"src": "[email protected]", "accept": ["grafana:443", "argocd:443"], "deny": ["kube-api:443"]},
{"src": "[email protected]", "accept": ["grafana:443"], "deny": ["argocd:443", "grafana:22"]},
],
}In the cluster: a ServiceAccount whose Role reaches only its own state Secret, a non-root userspace router, and a NetworkPolicy for the router pod:
apiVersion: v1
kind: ServiceAccount
metadata:
name: ts-router
namespace: tailscale
---
# The state Secret is created empty up front, so the Role never needs "create".
apiVersion: v1
kind: Secret
metadata:
name: ts-router-state
namespace: tailscale
type: Opaque
---
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: ts-router-state
namespace: tailscale
rules:
- apiGroups: [""]
resources: ["secrets"]
resourceNames: ["ts-router-state"] # its own state, nothing else
verbs: ["get", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
name: ts-router-state
namespace: tailscale
subjects:
- kind: ServiceAccount
name: ts-router
namespace: tailscale
roleRef:
apiGroup: rbac.authorization.k8s.io
kind: Role
name: ts-router-state
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: ts-router
namespace: tailscale
spec:
replicas: 1
strategy:
type: Recreate # one node identity per state Secret
selector:
matchLabels: {app: ts-router}
template:
metadata:
labels: {app: ts-router}
spec:
serviceAccountName: ts-router
securityContext:
runAsNonRoot: true
runAsUser: 1000
runAsGroup: 1000
seccompProfile: {type: RuntimeDefault}
containers:
- name: tailscale
image: tailscale/tailscale:v1.102.4 # pin; better, by digest
env:
- {name: TS_USERSPACE, value: "true"} # no TUN device, no NET_ADMIN
- {name: TS_KUBE_SECRET, value: ts-router-state}
- {name: TS_ROUTES, value: "10.96.0.0/16"}
- {name: TS_EXTRA_ARGS, value: "--login-server=https://hs.example.com"} # the tagged auth key sets the tag
- name: TS_AUTHKEY # single-use, tagged, minutes; synced from OpenBao
valueFrom:
secretKeyRef: {name: ts-router-authkey, key: authkey, optional: true}
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: [ALL]
resources:
requests: {cpu: 50m, memory: 64Mi}
limits: {memory: 256Mi}
volumeMounts:
- {name: tmp, mountPath: /tmp}
volumes:
- name: tmp
emptyDir: {}
---
# The router may reach the two services it fronts, DNS, and the tailnet
# (control server, relays, peers). Nothing else in the cluster.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: ts-router
namespace: tailscale
spec:
podSelector:
matchLabels: {app: ts-router}
policyTypes: [Ingress, Egress]
ingress: []
egress:
- to: # the API server, for its state Secret
- ipBlock: {cidr: 192.0.2.10/32} # your API server endpoint(s): kubectl get endpoints kubernetes
ports:
- {protocol: TCP, port: 6443}
- to:
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: kube-system}
podSelector:
matchLabels: {k8s-app: kube-dns}
ports:
- {protocol: UDP, port: 53}
- {protocol: TCP, port: 53}
- to:
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: monitoring}
podSelector:
matchLabels: {app.kubernetes.io/name: grafana}
ports:
- {protocol: TCP, port: 3000}
- to:
- namespaceSelector:
matchLabels: {kubernetes.io/metadata.name: argocd}
podSelector:
matchLabels: {app.kubernetes.io/name: argocd-server}
ports:
- {protocol: TCP, port: 8080}
- to:
- ipBlock:
cidr: 0.0.0.0/0
except: [10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16]
ports:
- {protocol: TCP, port: 443} # control server and DERP
- {protocol: UDP, port: 3478} # STUN
- {protocol: UDP, port: 41641} # peers on the default port; others relay via DERPThe router reads and writes its state Secret through the Kubernetes API, so
it needs egress to the API server; without that rule it never starts
tailscaled. On Cilium, the API server has its own identity and an
ipBlock for its address did not let the pod through in our test; use a
CiliumNetworkPolicy next to the NetworkPolicy:
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: ts-router-apiserver
namespace: tailscale
spec:
endpointSelector: {matchLabels: {app: ts-router}}
egress:
- toEntities: [kube-apiserver]
toPorts: [{ports: [{port: "6443", protocol: TCP}]}]With Headscale 0.29.4, do not add --advertise-tags=tag:k8s-router to a
node that joins with a tagged auth key: Headscale refuses the registration
with requested tags [tag:k8s-router] are invalid or not permitted. The key
already makes the node tag:k8s-router.
After the router has joined, delete the ts-router-authkey Secret. The node
key lives in ts-router-state from then on. With these variables,
the image's start program (containerboot, v1.102.4 source) runs tailscaled
with --tun=userspace-networking, --state=kube:ts-router-state and
--statedir=/tmp, and puts its socket at /tmp/tailscaled.sock. That is why
/tmp is the only writable mount. The Docker parameters page gives
/var/run/tailscale/tailscaled.sock as the socket default; if a later image
follows the page, set TS_SOCKET=/tmp/tailscaled.sock. Tailscale's Docker
parameters page recommends TS_AUTH_ONCE=true when state is persisted, so a
restart does not try to log in again.
Tailscale's policy file reference says an auto-approver for a route may also
advertise a subnet of it. Here that means tag:k8s-router could get
10.96.20.0/24 approved, never anything outside 10.96.0.0/16. Headscale's
routes page does not say either way; check with headscale nodes list-routes
after a change.
Prove it
From secure-tests/userspace-tailscale-subnet-router-kubernetes/. Headscale
0.29.4 runs in a container. Tailscale clients join with the same restrictions
as the pod above: uid 1000, --cap-drop ALL, a read-only root and userspace
networking. The router advertises the Service CIDR and, by mistake,
10.0.0.0/8. A laptop advertises the Service CIDR too:
k8s-router: joined as uid 1000, no capabilities
Some peers are advertising routes but --accept-routes is false
bob-laptop: joined as uid 1000, no capabilitiesOnly the exact route, from the tagged router, is approved:
headscale nodes list-routesk8s-router approved=10.96.0.0/16 available=10.0.0.0/8,10.96.0.0/16
bob-laptop approved= available=10.96.0.0/16The grants and their tests pass: ops reaches Grafana and Argo CD on 443, dev reaches Grafana only, and nobody reaches the API server through the router:
headscale policy check -f policy.hujsonPolicy is validThe manifests pass schema validation:
Summary: 6 resources found in 1 file - Valid: 6, Invalid: 0, Errors: 0, Skipped: 0In a lab cluster, with the Deployment, RBAC and policies above. Headscale
ran next to the cluster over HTTPS with a lab certificate (mounted into the
router and the laptops through SSL_CERT_FILE), and the NetworkPolicy got
two lab rules for Headscale and the laptops, which sat on a private network
standing in for the internet. Stand-ins answered at the Grafana
(10.96.20.5) and Argo CD (10.96.30.7) Service addresses. Two userspace
Tailscale clients joined as [email protected] (ops) and [email protected]
(dev), with --accept-routes.
What each person reaches through the router:
$ headscale nodes list-routes
ts-router-86595576ff-zlnxh tags=tag:k8s-router approved=10.96.0.0/16 available=10.96.0.0/16
alice (ops) grafana: 200 argocd: 200 kube-api: timeout
bob (dev) grafana: 200 argocd: timeout kube-api: timeoutWhat went wrong on the way:
NetworkPolicy without the API server rule:
boot: error setting up for running on Kubernetes: getting Tailscale state Secret ts-router-state:
Get "https://kubernetes.default.svc/api/v1/namespaces/tailscale/secrets/ts-router-state": context deadline exceeded
tagged auth key plus --advertise-tags=tag:k8s-router:
Received error: handling register with auth key: creating new node: requested tags [tag:k8s-router] are invalid or not permittedThe router pod:
$ kubectl -n tailscale exec deploy/ts-router -- id
uid=1000 gid=1000 groups=1000
$ kubectl auth can-i list secrets -n tailscale --as=system:serviceaccount:tailscale:ts-router
no
$ kubectl auth can-i get secret/ts-router-state -n tailscale --as=system:serviceaccount:tailscale:ts-router
yes
$ kubectl auth can-i get secret/ts-router-authkey -n tailscale --as=system:serviceaccount:tailscale:ts-router
no
router -> grafana (in the policy): 200
router -> cert-manager webhook (not in the policy): timed outAfter ts-router-authkey was deleted and the pod restarted, the router came
back as the same node from ts-router-state.
Mistakes people make
No egress to the API server
The router keeps its identity in a Secret and needs the API server for it. A NetworkPolicy without that rule leaves the pod running and the tailnet empty.
Privileged pods for a router that does not need them
Userspace mode needs no TUN device and no capabilities, and it is the image's
default (TS_USERSPACE defaults to true). The test ran the client as uid 1000
with every capability dropped.
Auto-approving a range for anyone
An auto-approver for 10.0.0.0/8 owned by a group lets any member's laptop
become a router for your whole private space. Approve exact CIDRs for a tag
that only ops can assign.
Granting the whole Service CIDR
A grant to 10.96.0.0/16 includes the API server, DNS and every internal
service. Name the services and ports.
Relying on NetworkPolicy to tell users apart
With SNAT, all tailnet traffic comes from the router pod. The cluster can restrict the router; only the tailnet policy can restrict people.
A Role with create and list on Secrets
The state Secret can be created ahead of time. Then get, update and patch
with resourceNames are enough.
A reusable auth key in a Secret forever
Use a single-use, tagged key that expires in minutes, and delete it once the router has joined.
Checklist
- Set
TS_USERSPACE=true; drop all capabilities; run as non-root with a read-only root. - Tag the router and restrict who may assign the tag.
- Auto-approve only the Service CIDR (subnets of it are approved too), only for the router's tag.
- Grant groups access to named service IPs and ports, never the whole CIDR.
- Add policy tests that deny the API server and other internal services.
- Pre-create the state Secret; give the Role
get,update,patchon it by name. - Use a single-use, short-lived, tagged auth key and delete it after the first join.
- Limit the router's egress with a NetworkPolicy.
A subnet router is a door into the cluster. Make it a narrow one with a list of rooms, not a hole in the wall.
H2-CPQE
Learn it on a live range
Private access and tailnets, in Edge and Post-Quantum Networking: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on identity and access
Self-hosted identity with Zitadel, private access with Headscale, break-glass and offboarding.
All identity and access guides