Disk encryption on Talos with LUKS2 and TPM, the secure way
The patch applied cleanly and the compliance sheet already says "encrypted at rest". The disk does not agree: it was formatted long ago, Talos cannot encrypt it in place, and the next reboot is when you find out.
The short answer
Encrypt both STATE and EPHEMERAL with LUKS2. On hardware with SecureBoot, seal the key to the TPM and refuse enrollment when SecureBoot is off. On cloud VMs without a TPM, use a network KMS, not nodeID or static keys. Lock EPHEMERAL to STATE, keep a second key slot for recovery, and verify every volume reports luks2.
On this page
What goes wrong
A Talos node keeps two sensitive volumes. STATE holds the machine
configuration: CA keys, tokens, the etcd encryption key. EPHEMERAL is
/var: etcd data on control plane nodes, container images, emptyDir
volumes and logs. Anyone who gets a copy of the disk (a snapshot, a returned
drive, a volume reattached to another VM) reads both.
Talos can encrypt both with LUKS2, but four details decide whether that means anything:
- Encryption happens only on empty volumes. There is no in-place
encryption. Adding the patch to a node that already has a formatted
EPHEMERALchanges nothing while the volume is mounted. The next time Talos prepares the volume (at boot), the release-1.14 source refuses a partition that already holds a filesystem:volumestatusshows phasefailedwith ablock dev type mismatcherror, and/varis not mounted. - TPM without SecureBoot is weak. The docs rate the TPM key strong only when used with SecureBoot. The key is bound to PCR 7, which records the SecureBoot state and enrolled keys, and to a signed policy on PCR 11. On Talos 1.14.1 a plain (non-SecureBoot) image cannot enroll a TPM key at all; the risk is a SecureBoot image booted with SecureBoot turned off in the firmware, where PCR 7 records "off".
- A static key on STATE is not a secret. Talos stores the STATE
encryption settings in cleartext in the
METApartition, next to the data. - nodeID is derived from the machine. The key comes from the node UUID and the partition label. It stops someone who walks away with the drive. It does not stop someone who knows the UUID, which is not a secret.
What the docs say
tpm- encrypt with the key derived from the TPM (strong, when used with SecureBoot).
Source: Sidero Labs docs, Disk Encryption
The
STATEvolume encryption configuration will be stored cleartext inMETAvolume, so it is not secure to usestatickeys forSTATEvolume.
Source: Sidero Labs docs, Disk Encryption
There is no in-place encryption support for the partitions right now, so to avoid losing data only empty partitions can be encrypted.
Source: Sidero Labs docs, Disk Encryption
Check that Secureboot is enabled in the EFI firmware. If Secureboot is not enabled, the enrollment of the key will fail.
Source: Sidero Labs docs, VolumeConfig
The guide's first example uses nodeID for both volumes, and the
checkSecurebootStatusOnEnroll switch appears only in the reference, not in
the guide or the SecureBoot page. So the copy-paste path gives you a key kind
the same page calls weak, and a TPM option that does not insist on
SecureBoot.
The secure configuration
Bare metal with SecureBoot and a TPM. Boot a SecureBoot image (from the Image Factory or your own signing key), then:
# talos-encryption-tpm.yaml
# For machines booted with SecureBoot. Apply before the volumes are formatted.
apiVersion: v1alpha1
kind: VolumeConfig
name: STATE
encryption:
provider: luks2
keys:
- slot: 0
tpm:
checkSecurebootStatusOnEnroll: true # refuse to enroll if SecureBoot is off
options:
pcrs: [7] # SecureBoot state (the default, written out)
- slot: 1
kms:
endpoint: https://kms.example.net:4443 # recovery path if the board or TPM is replaced
---
apiVersion: v1alpha1
kind: VolumeConfig
name: EPHEMERAL
encryption:
provider: luks2
keys:
- slot: 0
tpm:
checkSecurebootStatusOnEnroll: true
lockToState: true # cannot be unlocked once STATE is wiped or replaced
- slot: 1
kms:
endpoint: https://kms.example.net:4443
lockToState: trueTalos's own tooling agrees: talosctl cluster create qemu --presets iso-secureboot generates exactly the slot 0 above for both volumes. If your
base config already carries these keys, a patch adds to the key list instead
of replacing it, and Talos rejects the result with
VolumeConfig/STATE: duplicate key slot 0. Patch in only the slots you add.
Cloud VMs without a TPM or SecureBoot (for example, providers that boot
custom images with BIOS). Use a network KMS as the only key. The node must
reach the KMS before STATE unlocks, so it gets its network settings from
DHCP or the platform, not from the machine config. For the same reason the
KMS server certificate cannot rely on a custom CA from the machine config:
the docs say those are stored in STATE as well. Talos checks an
https:// KMS endpoint against the system roots; grpc:// means plaintext,
which sends the disk key across the network in clear.
# talos-encryption-kms.yaml
# For VMs without TPM/SecureBoot: the key is sealed by a KMS the disk never sees.
apiVersion: v1alpha1
kind: VolumeConfig
name: STATE
encryption:
provider: luks2
keys:
- slot: 0
kms:
endpoint: https://kms.example.net:4443 # reachable over the private network only
---
apiVersion: v1alpha1
kind: VolumeConfig
name: EPHEMERAL
encryption:
provider: luks2
keys:
- slot: 0
kms:
endpoint: https://kms.example.net:4443
lockToState: trueThe KMS server implements the Talos KMS API (Omni includes one; the
siderolabs/kms-client repository defines the API). Run it outside the
cluster it protects, restrict it to the node network, and back it up: without
it, the disks do not unlock.
Encrypting an existing node. Because only empty volumes are encrypted,
wipe them. Wiping STATE also makes every lockToState volume unreadable,
so when STATE is part of the change, wipe both in one step:
# EPHEMERAL only (STATE is already encrypted): stage, then wipe and reboot.
talosctl -n 10.10.1.21 apply-config -f worker.yaml --mode=staged
talosctl -n 10.10.1.21 reset --system-labels-to-wipe EPHEMERAL --reboot=true
# STATE and EPHEMERAL: wipe both; the node enters maintenance mode and loses
# its config, so apply it again. The console log line "server certificate
# issued" prints the fingerprint to pin.
talosctl -n 10.10.1.21 reset --system-labels-to-wipe STATE --system-labels-to-wipe EPHEMERAL --reboot=true
talosctl apply-config --insecure -n 10.10.1.21 --cert-fingerprint '<fingerprint>' -f worker.yamlDo control plane nodes one at a time, and only while etcd has quorum without the node you are wiping.
Prove it
Run on Talos 1.14.1 VMs in QEMU with a software TPM (swtpm). The KMS was the
reference kms-server from siderolabs/kms-client, built from source,
behind HTTPS with a certificate from a lab CA.
1. SecureBoot with the TPM and KMS configuration. A cluster from the SecureBoot ISO; the worker got the configuration above:
$ talosctl -n 10.10.3.3 get securitystate -o yaml
spec:
secureBoot: true
pcrSigningKeyFingerprint: 9C:42:05:91:48:E1:57:A0:...:3F:DC:03:2F
bootedWithUKI: true
$ talosctl -n 10.10.3.3 get volumestatus STATE -o yaml
phase: ready
encryptionProvider: luks2
encryptionFailedSyncs:
- 'error adding key slot 1 *keys.KMSKeyHandler: failed to seal KMS passphrase,
slot 1: rpc error: code = Unavailable desc = connection error: desc = "transport:
authentication handshake failed: tls: failed to verify certificate: x509:
certificate signed by unknown authority"'
configuredEncryptionKeys:
- kms
- tpm
encryptionSlot: 0
...
pcrs:
- 7
pubKeyPcrs:
- 11(Trimmed.) EPHEMERAL looked the same, with encryptionLockedToState: true.
The node booted and served with the TPM slot, and the recovery slot was
simply missing: the only sign was encryptionFailedSyncs. With the lab CA
added as a TrustedRootsConfig and a reboot:
STATE: {"phase":"ready","slot":0,"failed":["error adding key slot 1 *keys.KMSKeyHandler: failed to seal KMS passphrase, slot 1: rpc error: code = Unavailable desc ="]}
EPHEMERAL: {"phase":"ready","slot":0,"failed":[]}EPHEMERAL is opened after STATE, when the machine config and its CA are available; STATE never is. For a STATE recovery key, the KMS certificate must chain to a public root.
2. SecureBoot off. A cluster from the plain ISO, with a TPM:
$ talosctl -n 10.10.4.3 get volumestatus STATE -o json | jq -r .spec.errorMessage
error formatting and encrypting volume: no handlers available to get encryption keys from: 1 error occurred:
* failed to enroll the TPM2 key, as SecureBoot is disabled (and checkSecurebootOnEnroll is enabled)The volume stays failed and the node does not come up. Without
checkSecurebootStatusOnEnroll, on that plain image:
failed to calculate sealing policy digest: failed to calculate policy authorize: failed to read pcr signing public key: open /.extra/tpm2-pcr-public-key.pem: no such file or directory3. Encrypting a node that is already running. An EPHEMERAL encryption patch on an unencrypted node:
before: {"phase":"ready","provider":null}
patched, no reboot: {"phase":"ready","provider":null}
after reboot: {"phase":"failed","provider":null,"err":"block dev type mismatch: xfs != luks"}/var was not mounted. Only a wipe
(talosctl reset --system-labels-to-wipe EPHEMERAL) lets Talos format the
volume encrypted, with a key it can enroll.
4. The disk is useless elsewhere. The worker's disk image from step 1, read on the host:
$ blkid -p -o export /dev/loop0p3 ...
loop0p1: LABEL=EFI TYPE=vfat PART_ENTRY_NAME=EFI
loop0p2: PART_ENTRY_NAME=META
loop0p3: TYPE=crypto_LUKS PART_ENTRY_NAME=STATE
loop0p4: TYPE=crypto_LUKS PART_ENTRY_NAME=EPHEMERAL
$ cryptsetup luksDump /dev/loop0p4
Version: 2
Keyslots:
0: luks2
Cipher: aes-xts-plain64
PBKDF: argon2id
1: luks2
Cipher: aes-xts-plain64
PBKDF: argon2id
Tokens:
0: talos-tpm2Run the same checks on your nodes:
talosctl -n 10.10.1.21 get volumestatus STATE -o yaml
talosctl -n 10.10.1.21 get volumestatus EPHEMERAL -o yaml
talosctl -n 10.10.1.21 get securitystate -o yamlWhat you should see: phase: ready and encryptionProvider: luks2 on both
volumes, every configured slot present with no encryptionFailedSyncs,
encryptionLockedToState: true on EPHEMERAL, and secureBoot: true.
Mistakes people make
Ignoring encryptionFailedSyncs
A recovery key slot that cannot enroll does not stop the node. It boots on
the other slot, and the missing recovery path shows only in
encryptionFailedSyncs. Alert on it.
Patching a running node and trusting the patch
The patch only describes what Talos does to an empty volume. On a volume
formatted before the patch, it does nothing until the next boot, and then the
volume fails. Check encryptionProvider and phase in volumestatus, and
wipe and re-provision volumes that were formatted before the patch.
Sealing to a TPM with SecureBoot off
PCR 7 then records "SecureBoot disabled", and the firmware does not check
what boots. The docs call the TPM key strong only with SecureBoot, and
checkSecurebootStatusOnEnroll defaults to false. Set it to true so
enrollment fails loudly on such machines.
A static passphrase on STATE
Talos keeps the STATE encryption settings in cleartext in META. A static
key there is written next to the lock. Static keys are acceptable only on
other volumes, and only while STATE is encrypted.
Calling nodeID encryption "at rest encryption"
nodeID protects a drive that leaves the machine. In the release-1.14
source the passphrase is the node UUID followed by the partition label. The
UUID shows up in cloud APIs, inventories and talosctl output, so anyone who
has it and a copy of the disk can open the disk. Use it only as a second
slot, never as the only protection for secrets.
One key slot
Replace the motherboard or clear the TPM, and a TPM-only volume is gone. Keep a second slot (KMS) and test it before you need it. When rotating, keep one working key at all times and reboot between steps.
Checklist
STATEandEPHEMERALboth have aVolumeConfigwithprovider: luks2.- TPM keys set
checkSecurebootStatusOnEnroll: true, andsecuritystateshowssecureBoot: true. - Machines without TPM or SecureBoot use a KMS key, not
nodeIDorstatic, for STATE. - No
statickey is used on STATE. EPHEMERALand user volumes uselockToState: true.- Each volume has a second key slot for recovery, and recovery was tested.
- The KMS runs outside the cluster, on the private network, with backups.
talosctl get volumestatusshowsencryptionProvider: luks2on every node.- Nodes provisioned before the patch were wiped and re-provisioned one at a time.
Disk encryption is the rare control that fails silently in both directions: it can be configured and absent, or present and unlockable by anyone. Check the volume, not the YAML.
H2-CSPE
Learn it on a live range
Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on nodes and clusters
Talos Linux, Kubernetes API hardening, service account tokens, RBAC and the cloud underneath.
All nodes and clusters guides