Kubernetes and Talos API access over a tailnet, the secure way
Your API server has a public address because kubectl needs to reach it from your laptop, from CI and from the new contractor. So does every botnet, and the botnet never takes a day off.
The short answer
Run the Tailscale system extension on control plane nodes with a one-off, tagged auth key. Allow ports 6443 and 50000 only from the tailnet range and the node network, add the tailnet name to both API certificates, and use tailnet grants so only an admin group reaches those ports. Remove all public ingress to them at the cloud firewall.
On this page
What goes wrong
Two APIs control a Talos cluster. The Kubernetes API on 6443 runs every workload. The Talos API on 50000 reboots, resets and wipes nodes. Both use TLS client certificates or tokens, and both are commonly published on a public address so that people and pipelines can reach them.
A public API server is found by scanners within hours. Authentication holds until one of these happens: a leaked kubeconfig, a token with a long life in a CI log, a bug in the API server's pre-authentication path, or an anonymous endpoint someone enabled for a health check. Each of those is only exploitable if the attacker can connect.
The common fix, a VPN, often brings its own problem: one shared key, a flat network, and every VPN user able to reach every port on every node.
A tailnet can do better, but the default setup from the Talos extension README uses a plain auth key inside the machine config, and its subnet router example publishes every Kubernetes Service to the whole tailnet.
What the docs say
Adds https://tailscale.com network interfaces as system extensions. This means you can access your talos nodes from machines you have configured with tailscale
Source: Sidero Labs extensions, Tailscale README
Be very careful with reusable keys! These can be very dangerous if stolen.
Source: Tailscale docs, Auth keys
Avoid embedding Talos API credentials in automation unless you can properly scope and restrict their permissions.
Source: Sidero Labs docs, Talos Security Checklist
apidand Kubernetes API are wide open
Source: Sidero Labs docs, Ingress Firewall (recommended control plane rules)
The Talos docs have no guide for private API access, and the recommended
control plane rules allow 50000 and 6443 from 0.0.0.0/0 and ::/0. The extension README
puts TS_AUTHKEY in the machine config, where anyone who can read the config
reads the key, and it does not say which kind of key to use. Its subnet
routing example advertises the whole service range, 10.96.0.0/12, to the
tailnet.
The secure configuration
1. Build the node image with the Tailscale extension. Add it to your Image Factory schematic next to your other extensions:
# schematic.yaml
customization:
systemExtensions:
officialExtensions:
- siderolabs/tailscale2. Create one auth key per node. In the Tailscale admin console (or
Headscale), create a key that is one-off (not reusable), pre-approved,
tagged tag:talos-cp, and expires in one day. Once the node has joined, the
key is spent, so a copy in the machine config is worth nothing.
3. Patch the control plane nodes. In block mode the Talos ingress
firewall drops everything that no rule allows, so the patch carries the
kubelet, trustd, etcd and CNI rules from the Talos Ingress Firewall guide as
well as the two API rules. The example node network is 10.10.0.0/16 with
control plane nodes at 10.10.1.11 to 10.10.1.13, and the CNI is Cilium
with VXLAN on UDP 8472.
# talos-tailnet-controlplane.yaml
# talosctl gen config ... --config-patch-control-plane @talos-tailnet-controlplane.yaml
apiVersion: v1alpha1
kind: ExtensionServiceConfig
name: tailscale
environment:
- TS_AUTHKEY=tskey-auth-kExample1CNTRL-ExampleOneOffKey # one-off, tagged, 1-day expiry
- TS_AUTH_ONCE=true # log in once; restarts reuse the saved state, not the spent key
- TS_USERSPACE=false # already the extension's setting (kernel networking, a tailscale0 interface); explicit here
- TS_HOSTNAME=cp-1
---
apiVersion: v1alpha1
kind: NetworkDefaultActionConfig
ingress: block
---
apiVersion: v1alpha1
kind: NetworkRuleConfig
name: apid-ingress
portSelector:
ports:
- 50000
protocol: tcp
ingress:
- subnet: 100.64.0.0/10 # tailnet IPv4 addresses; tailnet grants narrow this to admins
- subnet: fd7a:115c:a1e0::/48 # tailnet IPv6 addresses (MagicDNS names resolve to both)
- subnet: 10.10.0.0/16 # node network (control plane proxies to workers)
---
apiVersion: v1alpha1
kind: NetworkRuleConfig
name: kubernetes-api-ingress
portSelector:
ports:
- 6443
protocol: tcp
ingress:
- subnet: 100.64.0.0/10
- subnet: fd7a:115c:a1e0::/48
- subnet: 10.10.0.0/16
---
apiVersion: v1alpha1
kind: NetworkRuleConfig
name: kubelet-ingress
portSelector:
ports:
- 10250
protocol: tcp
ingress:
- subnet: 10.10.0.0/16
---
apiVersion: v1alpha1
kind: NetworkRuleConfig
name: trustd-ingress
portSelector:
ports:
- 50001
protocol: tcp
ingress:
- subnet: 10.10.0.0/16
---
apiVersion: v1alpha1
kind: NetworkRuleConfig
name: etcd-ingress
portSelector:
ports:
- 2379-2380
protocol: tcp
ingress:
- subnet: 10.10.1.11/32
- subnet: 10.10.1.12/32
- subnet: 10.10.1.13/32
---
apiVersion: v1alpha1
kind: NetworkRuleConfig
name: cni-vxlan
portSelector:
ports:
- 8472
protocol: udp
ingress:
- subnet: 10.10.0.0/16
---
apiVersion: v1alpha1
kind: KubeAPIServerConfig
certExtraSANs:
- cp-1.example-tailnet.ts.net # the MagicDNS name kubectl will use
---
machine:
certSANs:
- cp-1.example-tailnet.ts.net # the same name for the Talos API certificateApply firewall changes to a running node with
talosctl apply-config --mode=try first: the Talos docs warn that a wrong
ingress configuration can make the node unreachable over the Talos API.
Workers need the same firewall default and the worker rules from
Talos Linux hardening. They do not need
the tailnet unless you manage them directly.
The Tailscale extension README calls userspace networking "the default".
That is the default of containerboot, the program the extension runs; the
extension's own service definition sets TS_USERSPACE=false, so the node
gets a kernel tailscale0 interface and the host firewall sees tailnet
source addresses without any setting from you.
4. Let only admins reach the ports. In the tailnet policy file, only
group:platform-admins reaches 6443 and 50000, and a read-only group reaches
6443 only. What each person can do once connected is still Kubernetes RBAC
and the Talos role in their certificate. The file below uses Tailscale's
grants syntax; if you run Headscale, check which policy fields your version
supports before you rely on it.
{
"groups": {
"group:platform-admins": ["[email protected]", "[email protected]"],
"group:platform-readers": ["[email protected]"]
},
"tagOwners": {
"tag:talos-cp": ["group:platform-admins"]
},
"grants": [
{
"src": ["group:platform-admins"],
"dst": ["tag:talos-cp"],
"ip": ["tcp:6443", "tcp:50000"]
},
{
"src": ["group:platform-readers"],
"dst": ["tag:talos-cp"],
"ip": ["tcp:6443"]
}
]
}5. Point clients at the tailnet name and close the public door.
kubectl config set-cluster demo --server=https://cp-1.example-tailnet.ts.net:6443
talosctl config endpoint cp-1.example-tailnet.ts.net
talosctl config node 100.101.102.103 # the node's tailnet IP (tailscale status), not its nameUse the name for the endpoint, which your machine resolves through MagicDNS.
Use an address for node: Talos resolves node names on the node itself, and
the node does not use the tailnet's DNS. With the name as the node,
talosctl fails with name resolver error: produced zero addresses.
If you run Headscale, Headscale 0.29.4 accepts this policy file
(headscale policy check: Policy is valid) and enforced it in our test.
Then remove every public inbound rule for 6443 and 50000 from the cloud
firewall and delete any public load balancer in front of the API. Host
components on the nodes reach the API server through KubePrism on
localhost:7445, which balances across the cluster endpoint, the local API
server on control plane nodes, and every control plane address found by
Cluster Discovery. The cluster endpoint itself (endpoint in
KubeClusterConfig) is still in the machine config: if it names the public
load balancer, move it to a private address, such as a DNS name for the
control plane nodes on the node network, before you delete the balancer.
Prove it
Run on a Talos 1.14.1 control plane in QEMU (node address 10.10.5.2),
installed from a schematic with siderolabs/tailscale and the patch above.
The tailnet was Headscale 0.29.4 with the policy above; for the lab, the
patch added TS_EXTRA_ARGS=--login-server=... and the Headscale CA as a
TrustedRootsConfig, the etcd rule named the one control plane node, and
the tailnet name was cp-1.tail.example.com. alice and bob were in
group:platform-admins, carol in group:platform-readers.
1. The node joined the tailnet with its tag:
$ headscale nodes list
ID | Hostname | Name | ... | User | Tags | IP addresses | ... | Connected
1 | cp-1 | cp-1 | ... | tagged-devices | tag:talos-cp | 100.64.0.1 | ... | online
$ talosctl -n 10.10.5.2 service ext-tailscale
STATE Running2. The APIs answer on the tailnet name (bob, a client with a kernel
tailscale0 interface, using the configs from step 5):
$ kubectl get --raw /readyz
ok
$ talosctl version # config node = cp-1.tail.example.com
error getting version from node cp-1.tail.example.com: rpc error: code = Unavailable desc = name resolver error: produced zero addresses
$ talosctl -n 100.64.0.1 version
Tag: v1.14.1No certificate name errors: certExtraSANs and certSANs carry the tailnet
name. The first talosctl call shows why node must be an address.
3. The node's address is closed off the tailnet (from a host outside the node network and the tailnet):
203.0.113.9 -> 10.10.5.2:6443: closed
203.0.113.9 -> 10.10.5.2:50000: closed4. A reader cannot reach the Talos API:
alice (platform-admins) 6443: open 50000: open
carol (platform-readers) 6443: open 50000: blocked5. A spent key does not matter after the first join. The one-off key
showed used: true; after talosctl service ext-tailscale restart and after
a reboot the node came back online as the same tailnet node (id=1), from
its saved state.
Mistakes people make
The tailnet name as the talosctl node
talosctl config node <name> makes the node resolve its own name, without
tailnet DNS. Keep the name for the endpoint and use the tailnet IP for the
node.
Putting a reusable key in the machine config
The machine config is readable by anyone with the admin talosconfig, by anyone with access to the repository it lives in, and, when it is delivered as cloud user data, by anything on the node that can reach the metadata service. A reusable key there lets anyone add devices to your tailnet. Use a one-off, tagged, short-lived key per node.
Advertising the service range
TS_ROUTES=10.96.0.0/12 makes every ClusterIP Service reachable from the
tailnet, including ones that never had authentication because they were
"internal". If you need one Service, expose that Service, and grant access to
it by tag.
Opening the firewall to 100.64.0.0/10 and stopping there
The CGNAT range (and the tailnet IPv6 prefix) covers every device in the tailnet, including phones and shared nodes. The host firewall rule is the outer wall; the tailnet grants decide who gets in. Write both.
Forgetting the certificate names
Without the tailnet name in certExtraSANs and machine.certSANs, kubectl
and talosctl refuse the connection, and people "fix" it with
insecure-skip-tls-verify. Add the names before you switch endpoints.
Keeping the public load balancer "just in case"
A public load balancer in front of 6443 keeps the old exposure alive. Keep a documented break-glass path instead, such as a bastion in the VPC, and test it.
Checklist
- Control plane nodes run the
siderolabs/tailscaleextension from your schematic. - Each node joined with a one-off, pre-approved, tagged auth key that has expired.
TS_AUTH_ONCE=trueis set, and nothing setsTS_USERSPACE=true(the extension runs kernel networking by default).- The host firewall is in block mode; 6443 and 50000 allow only the tailnet IPv4 and IPv6 ranges and the node network.
- The kubelet, trustd, etcd and CNI rules are in place on control plane nodes, and the change was applied with
--mode=tryfirst. - The tailnet policy grants 6443 and 50000 to the admin group only.
certExtraSANsandmachine.certSANscontain the tailnet names.- Kubeconfig and talosconfig point at tailnet names.
- The cluster endpoint in
KubeClusterConfignames a private address. - The cloud firewall has no public inbound rule for 6443 or 50000, and no public API load balancer exists.
- No
TS_ROUTESadvertises the pod or service ranges.
The best API server exploit is the one that never gets a TCP handshake. Make the tailnet the only road in, and put a gate on the road.
H2-CSPE
Learn it on a live range
Cluster networking and policy, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on nodes and clusters
Talos Linux, Kubernetes API hardening, service account tokens, RBAC and the cloud underneath.
All nodes and clusters guides