Talos upgrades that keep your hardening, the secure way
The upgrade finished, every node says Ready, and the changelog promised new isolation defaults. Two days later the gVisor pods will not start, and the isolation you read about was never switched on for your cluster.
The short answer
Upgrade one minor version at a time, always with an explicit installer: your platform's installer and the schematic the node was built from. Snapshot extensions and the kernel command line before, compare after. Then add the security documents that new clusters get by default, because an upgrade keeps your old configuration exactly as it was.
On this page
What goes wrong
A Talos upgrade replaces the operating system image. It does not replace your machine configuration. Three things follow from that, and none of them shows up as an error.
The installer decides what you get. The upgrade call names an installer
image, and that image carries the system extensions (gVisor, firmware,
drivers). talosctl upgrade without --image uses the metal installer with
the empty schematic, for the version of talosctl you run. On a node that was
built with the gVisor extension, that upgrade leaves gVisor out. The runsc
handler disappears, and every pod that asks for it fails to start.
Kernel arguments are a moving target. The Talos docs do not agree on
where kernel arguments come from in an upgrade. The Bootloader page says that
on GRUB machines they are taken from the installer image, built from a
schematic's customization.extraKernelArgs. The Image Factory page says
installer images ignore kernel args. The Bootloader page also warns that
the upgrade API introduced in Talos 1.13 does not apply
.machine.install.extraKernelArgs when grubUseUKICmdline is false, which
is the default for installs that were upgraded to 1.12 from older versions.
Its fix is talosctl upgrade --legacy or a move to the UKI command line.
Either way, an init_on_free=1 you added can be there before the upgrade and
gone after, so check the node, not the docs.
New defaults stay in the release notes. Talos 1.14 can run containerd,
the kubelet and all pods in a dedicated PID and mount namespace, away from
machined (workload isolation). Configs generated by Talos 1.14 turn it on, and
also carry a VolumeConfig that mounts /var with nosuid,nodev. A
cluster upgraded from 1.13 keeps its 1.13 configuration: it has no
SecurityProfileConfig document, so isolation stays off. (On 1.13.10 in our
test, /var was already mounted nosuid,nodev without that document; add
it anyway, so the setting is written down.)
Talos also moves its own defaults. From 1.13.10 to 1.14.1, the kernel
command line lost init_on_alloc=1 and nvme_core.io_timeout=4294967295.
init_on_alloc did not go away: the 1.14.1 kernel is built with
CONFIG_INIT_ON_ALLOC_DEFAULT_ON=y. A check that compares the whole command
line before and after stops every correct upgrade.
What the docs say
This matters because it means Talos clusters start compliant and stay compliant across upgrades, without operator intervention.
Source: Sidero Labs docs, Talos Default Hardening and CIS Compliance
clusters upgraded from older versions do not have the document and keep the old (non-isolated) behavior unless it is added.
Source: Sidero Labs docs, SecurityProfileConfig
The new upgrade API introduced in Talos v1.13 does not apply
.machine.install.extraKernelArgswhengrubUseUKICmdlineis set tofalse.
Source: Sidero Labs docs, Bootloader
The upgrade process should handle such changes transparently, but this migration is only tested between adjacent minor releases.
Source: Sidero Labs docs, Upgrading Talos Linux
The first quote follows a table of CIS controls, most of which the page says
are enforced at the platform level. It does not cover new defaults that live
in configuration documents, and the second quote says so on a different page.
On kernel arguments the Bootloader page says GRUB takes them from the
installer image, while the Image Factory page says "installer and
initramfs images only support system extensions (kernel args and META are
ignored)". The upgrade guide does tell you to use the schematic ID the node
was installed from, but its only example is the metal installer with the
empty schematic.
The secure configuration
1. Keep the inputs of every node in git. The schematic file, its ID, the
installer image name for the node's platform (with -secureboot for
SecureBoot nodes), and the Talos version it runs.
talos/
schematic.yaml # extensions and extra kernel arguments
schematic.id # the ID Image Factory returned for schematic.yaml
installer # metal-installer, aws-installer, metal-installer-secureboot, ...
kernel-args # the kernel arguments you set, one per line (the script checks them)
patches/ # your machine config patches (see the git + age page)2. Upgrade with a script that pins the installer and compares the node before and after.
# talos-upgrade.sh
# Usage: bash talos-upgrade.sh <node-ip> <talos-version>, for example 10.10.1.11 v1.14.1
# Upgrade one minor version at a time (1.12 -> latest 1.13 -> 1.14).
set -euo pipefail
node="$1"
version="$2"
installer="$(cat talos/installer)"
schematic="$(cat talos/schematic.id)"
image="factory.talos.dev/${installer}/${schematic}:${version}"
snapshot() {
# $1 is a file prefix. Extension names (the "schematic" entry carries the
# schematic ID) and the kernel command line, one argument per line.
talosctl -n "$node" get extensions -o jsonpath='{.spec.metadata.name}{"\n"}' | sort > "$1-extensions.txt"
talosctl -n "$node" get cmdline -o jsonpath='{.spec.cmdline}' | tr ' ' '\n' | sort > "$1-cmdline.txt"
}
snapshot "before-${node}"
# GRUB nodes that still set .machine.install.extraKernelArgs with
# grubUseUKICmdline: false: the Bootloader docs say only --legacy applies them.
talosctl -n "$node" upgrade --image "$image" --wait
snapshot "after-${node}"
# Every extension present before must be present after.
missing="$(comm -23 "before-${node}-extensions.txt" "after-${node}-extensions.txt")"
if [ -n "$missing" ]; then
printf 'STOP: extensions lost on %s:\n%s\n' "$node" "$missing" >&2
exit 1
fi
# Every kernel argument you set (talos/kernel-args, one per line) must be present.
missing="$(sort talos/kernel-args | comm -23 - "after-${node}-cmdline.txt")"
if [ -n "$missing" ]; then
printf 'STOP: kernel arguments lost on %s:\n%s\n' "$node" "$missing" >&2
exit 1
fi
# Talos's own defaults change between versions: report, do not stop.
changed="$(comm -3 "before-${node}-cmdline.txt" "after-${node}-cmdline.txt" | tr -d '\t')"
[ -z "$changed" ] || printf 'note: default kernel arguments changed on %s (check the release notes):\n%s\n' "$node" "$changed"
echo "ok: ${node} runs ${version} with the same extensions and kernel arguments"Run it on one node, check workloads on that node, then continue node by node.
If the nodes run gVisor, the extension alone is not enough on Talos: gVisor
needs user namespaces, and Talos sets user.max_user_namespaces to 0. The
extension's README gives the override, which undoes a KSPP setting, so set
it on sandbox nodes only:
machine:
sysctls:
user.max_user_namespaces: "11255" # gVisor needs user namespaces; sandbox nodes onlyWithout it, gVisor pods fail with gofer: fork/exec /proc/self/exe: no space left on device.
3. After the last node is upgraded, add the new security defaults. For a cluster upgraded to 1.14 from an older version:
# talos-post-upgrade-1.14.yaml
# Apply after every node runs 1.14; Talos 1.13 has no SecurityProfileConfig kind.
apiVersion: v1alpha1
kind: SecurityProfileConfig
workloadIsolation: true # new 1.14 clusters get this; upgraded ones do not
---
apiVersion: v1alpha1
kind: VolumeConfig
name: EPHEMERAL
mount:
secure: true # nosuid, nodev on /var, as in a fresh 1.14 configtalosctl -n 10.10.1.11 patch machineconfig -p @talos-post-upgrade-1.14.yamlWorkload isolation moves the container plane into a new namespace, so roll it out like an upgrade: one node, check, then the rest. The in-tree iSCSI volume plugin stops working with it; use a CSI driver.
4. Find the next set of new defaults yourself. Before each minor upgrade, generate a throwaway config for the old contract and for the new one and compare the security settings. This needs no cluster:
for v in v1.13 v1.14; do
talosctl gen config demo https://203.0.113.10:6443 --talos-version "$v" \
--output-types controlplane --output "cp-$v.yaml"
done
grep -nE 'workloadIsolation|secure: true' cp-v1.13.yaml cp-v1.14.yamlKeep your own regenerated configs on the old --talos-version contract until
you decide to adopt the new defaults. Changing the contract changes the
generated config.
Prove it
Run on two Talos VMs (QEMU), a control plane at 10.10.2.2 and a worker at
10.10.2.3, installed with Talos 1.13.10 from a schematic with the gVisor
extension and init_on_free=1, and upgraded to 1.14.1. The worker had the
user.max_user_namespaces override above.
1. The new defaults are missing from an old contract:
$ grep -nE 'workloadIsolation|secure: true' cp-v1.13.yaml cp-v1.14.yaml
cp-v1.14.yaml:107: secure: true # Enable secure mount options (nosuid, nodev).
cp-v1.14.yaml:139:workloadIsolation: true # Enable workload isolation (run the container plane inside the sandbox namespace).After the upgrade, the running config of both nodes still had no
SecurityProfileConfig and no VolumeConfig document.
2. An upgrade without --image removes the extension:
$ talosctl -n 10.10.2.3 upgrade --wait
$ talosctl -n 10.10.2.3 get extensions
NODE NAMESPACE TYPE ID VERSION NAME VERSION
10.10.2.3 runtime ExtensionStatus 0 1 schematic 376567988ad370138ad8b2698212367b8edcb69b5fd68c80be1f2ec7d603b4ba
$ kubectl run gvt --image=alpine:3.22 --restart=Never --overrides='{"spec":{"runtimeClassName":"gvisor",...}}' -- dmesg
Failed to create pod sandbox: ... unable to get OCI runtime for sandbox "a24c951a...": no runtime for "runsc" is configured376567... is the empty schematic. init_on_free=1 was gone from the
command line too.
3. The script stops a wrong upgrade and passes a right one:
$ bash talos-upgrade.sh 10.10.2.3 v1.14.1 # talos/schematic.id set to the empty schematic
STOP: extensions lost on 10.10.2.3:
gvisor
(exit 1)
$ bash talos-upgrade.sh 10.10.2.3 v1.14.1 # the node's own schematic
note: default kernel arguments changed on 10.10.2.3 (check the release notes):
init_on_free=1
ok: 10.10.2.3 runs v1.14.1 with the same extensions and kernel arguments
$ kubectl logs gvt
[ 0.000000] Starting gVisor...(The note shows init_on_free=1 coming back, because this run repaired the
node from step 2.) On the control plane, 1.13.10 to 1.14.1 with the node's
schematic:
note: default kernel arguments changed on 10.10.2.2 (check the release notes):
init_on_alloc=1
nvme_core.io_timeout=4294967295
ok: 10.10.2.2 runs v1.14.1 with the same extensions and kernel arguments
$ talosctl -n 10.10.2.2 read /proc/config.gz | zcat | grep INIT_ON_ALLOC
CONFIG_INIT_ON_ALLOC_DEFAULT_ON=yThe first version of this script compared the whole command line and
stopped here with STOP: cmdline lost, although nothing you set was lost.
4. The schematic is the one in git:
$ talosctl -n 10.10.2.2,10.10.2.3 get extensions
10.10.2.2 gvisor 20260831.0
10.10.2.2 schematic ddee0e5af521d05159a24d92f5ebcb8512413933156ff6220991f517590b48a4
10.10.2.3 gvisor 20260831.0
10.10.2.3 schematic ddee0e5af521d05159a24d92f5ebcb8512413933156ff6220991f517590b48a45. Workload isolation after the patch:
$ talosctl -n 10.10.2.3 patch machineconfig -p @talos-post-upgrade-1.14.yaml
Applied configuration without a reboot
$ talosctl -n 10.10.2.3 get mc v1alpha1 -o yaml | grep -A1 'kind: SecurityProfileConfig'
kind: SecurityProfileConfig
workloadIsolation: trueA plain pod and a gVisor pod on that node both ran to completion afterwards.
Mistakes people make
Letting the default image win
talosctl upgrade without --image installs the metal installer with the
empty schematic for your talosctl version. That leaves out every extension
the node's schematic carried, and on a cloud or SecureBoot node it is the
wrong installer as well.
Skipping minor versions
Talos tests configuration migration only between adjacent minor releases. Jumping from 1.12 to 1.14 is a path nobody tested. Upgrade to the latest patch of each minor version in turn, as the upgrade guide recommends.
Reading "secure by default" as "secure after upgrade"
Controls enforced in code stay on. Defaults that arrive as new configuration documents only appear in new configs. Read the "Machine configuration changes" section of each upgrade guide and add what you want.
Bumping the contract by accident
Regenerating configs without --talos-version, or with a newer one, pulls in
every new default at once, together with changed fields. Do it on purpose,
review the diff, and apply it like any other change.
Upgrading every node at once
A bad installer reference on one node is a small outage. On all nodes at once it is a cluster rebuild. Upgrade one node, check its workloads, then continue.
Checklist
schematic.yaml,schematic.idand the installer name are in git.- Every upgrade passes
--image factory.talos.dev/<installer>/<schematic>:<version>. - GRUB nodes with
.machine.install.extraKernelArgsupgrade with--legacyor move to the UKI command line. - Upgrades go one minor version at a time, to the latest patch of each.
- Extensions and the kernel command line are captured before and compared after each node.
- One node is upgraded and checked before the others.
- After the last node, the security documents new clusters get by default are added.
workloadIsolation: trueand/varsecure mount options are present on 1.14 nodes.- The
--talos-versioncontract changes only as a reviewed change.
An upgrade does exactly what you tell it, and nothing you forgot to say. Say the installer out loud, and diff the node when it comes back.
H2-CSPE
Learn it on a live range
Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on nodes and clusters
Talos Linux, Kubernetes API hardening, service account tokens, RBAC and the cloud underneath.
All nodes and clusters guides