Runtime detection and observability
Tetragon runtime detection on Talos, the secure way
Talos took away SSH, the shell and the package manager, which is wonderful right up until you need to know what ran on a node last night. With no shell to look around in, the kernel's own view of every process is the evidence you get. Better make sure it is being recorded.
The short answer
Install Tetragon with Helm in kube-system with the Talos values from its docs. Turn on process credentials and namespaces, keep the host and kube-system events that the default deny list drops, redact secrets from arguments, keep gRPC on the Unix socket, watch setns from any container mount namespace, and ship the JSON export off the node.
On this page
What goes wrong
Talos Linux has no SSH, no shell and a read-only root filesystem. That removes most of the ways an attacker persists on a node. It also removes most of the ways a defender investigates one. You cannot log in and look.
Tetragon fills that gap. It watches process execution, file access, network calls and privilege changes from inside the kernel with eBPF, and it tags each event with the pod, namespace and container it came from.
Three defaults work against you. The Helm chart's export deny list drops events
from kube-system and from the host itself, which is exactly where an
attacker goes after a container escape. And the gRPC API can be exposed on a
TCP port, which lets anyone who reaches it read every event, or load
policies.
What the docs say
The following Helm values configuration is required to install Tetragon on Talos Linux:
Source: Tetragon docs, Deploy on Kubernetes, Requirements for Talos Linux
By default, pods in the kube-system namespace are filtered-out.
Source: Tetragon docs, Deploy on Kubernetes
WARNING: Exposing gRPC on a TCP socket without TLS client verification exposes Tetragon to unprivileged users on the host or with network access.
Source: Tetragon docs, Helm chart reference
The Talos requirement is a host mount for /sys/kernel/tracing. On Talos
1.14.1 with Tetragon 1.7.1, the agent started and kprobe policies fired
without it (see Prove it); it costs nothing, so keep it for the versions the
docs had in mind.
The note about kube-system is short and easy to skip. The chart's default
exportDenyList also contains an empty namespace, "", which means "not in
any pod": the host. On a Talos node, host processes are the Talos services
themselves, and anything an attacker starts after escaping a container.
The secure configuration
Helm values for Talos:
# values-talos.yaml: Tetragon Helm values for Talos Linux.
# From the Tetragon docs for Talos 1.12+. Not needed on Talos 1.14.1 in our test; harmless, keep it.
extraHostPathMounts:
- name: sys-kernel-tracing
mountPath: /sys/kernel/tracing
tetragon:
# Add process credentials and namespaces to every event (off by default).
enableProcessCred: true
enableProcessNs: true
# gRPC stays on the local Unix socket; never expose it on a TCP port.
grpc:
enabled: true
address: "unix:///var/run/tetragon/tetragon.sock"
# The default deny list drops events from the host ("") and kube-system.
# On Talos, the host is exactly where an attacker would go next: keep it.
exportDenyList: |-
{"health_check":true}
# Mask secrets passed on command lines before events leave the node.
redactionFilters: |-
{"redact": ["(?:--password|--token|--secret)(?:\\s+|=)(\\S*)"]}
prometheus:
enabled: truehelm repo add cilium https://helm.cilium.io
helm repo update
helm install tetragon cilium/tetragon -n kube-system --version 1.7.1 -f values-talos.yaml
kubectl rollout status -n kube-system ds/tetragon -wA first detection that matters on every cluster: a process in a container
calling setns() to join another namespace, the core move of nsenter-style
escapes. Select on the mount namespace, not the PID namespace: a pod with
hostPID: true, the usual starting point for nsenter -t 1, shares the
host's PID namespace but still has its own mount namespace. runc joins
container namespaces on every kubectl exec; leave it out by binary.
# Container escape attempts: setns() from any process outside the host mount namespace.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
name: namespace-access
spec:
kprobes:
- call: "sys_setns"
syscall: true
args:
- index: 0
type: "int" # fd of the target namespace
- index: 1
type: "int" # namespace type flag (CLONE_NEWNS, CLONE_NEWNET, ...)
selectors:
- matchNamespaces:
- namespace: Mnt # hostPID pods share the host PID namespace, never its mount namespace
operator: NotIn
values:
- "host_ns"
matchBinaries:
- operator: NotIn
values:
- "/usr/bin/runc" # runc joins container namespaces on every kubectl execIn an ordinary pod this policy stays quiet even during an attempt. Talos runs
pods with the RuntimeDefault seccomp profile, and its default Pod Security
level (baseline) forbids Unconfined, so the setns() call is refused
before the kernel function runs. Tetragon still records the nsenter
process and its failed exit. The policy fires for pods that could actually
escape: privileged ones.
kubectl apply -f namespace-access.yamlShip the export off the node. The chart writes JSON events to a file that the
export-stdout container prints; collect that container's logs with your log
agent (Alloy, Fluent Bit) into Loki or your SIEM. On Talos, Tetragon's events
are the process history of the node: keep them for as long as you keep other
security logs.
Prove it
Run on two Talos 1.14.1 VMs (a control plane and a worker) with Tetragon 1.7.1 installed from the chart with the values above.
1. The agent runs on every node:
$ kubectl get pods -n kube-system -l app.kubernetes.io/name=tetragon -o wide
NAME READY STATUS RESTARTS NODE
tetragon-nrwjx 2/2 Running 0 harden-worker-1
tetragon-xxkgb 2/2 Running 0 harden-controlplane-1
$ kubectl exec -n kube-system ds/tetragon -c tetragon -- tetra status
Health Status: runningWithout extraHostPathMounts the same: 2/2 Running, no tracing host path in
the DaemonSet, /sys/kernel/tracing readable in the pod, and the policy
below fired for the escape (5 events).
2. An escape attempt from an ordinary pod (team-a/debug, alpine;
tetra getevents -o compact output below, with its icons removed):
kubectl exec -n kube-system ds/tetragon -c tetragon -- tetra getevents -o compact --namespace team-a
kubectl exec -n team-a debug -- nsenter -t 1 -m -u -n -i -p -- trueprocess team-a/debug /usr/bin/nsenter "-t 1 -m -u -n -i -p -- true"
exit team-a/debug /usr/bin/nsenter "-t 1 -m -u -n -i -p -- true" 1No setns event: the pod has Seccomp: 2 in /proc/self/status, and a
direct os.setns() from Python in such a pod fails with [Errno 1] Operation not permitted without reaching the kprobe. A pod asking for
Unconfined is refused by admission:
violates PodSecurity "baseline:latest": seccompProfile (pod must not set securityContext.seccompProfile.type to "Unconfined").
3. A real escape (a privileged pod with hostPID: true in a namespace
allowed to run it):
process escape-test/priv /usr/bin/nsenter "-t 1 -m -u -n -i -p -- cat /etc/os-release" CAP_SYS_ADMIN
setns escape-test/priv /usr/bin/nsenter ipc CAP_SYS_ADMIN
setns escape-test/priv /usr/bin/nsenter uts CAP_SYS_ADMIN
setns escape-test/priv /usr/bin/nsenter net CAP_SYS_ADMIN
setns escape-test/priv /usr/bin/nsenter pid CAP_SYS_ADMIN
setns escape-test/priv /usr/bin/nsenter mnt CAP_SYS_ADMIN
exit escape-test/priv /usr/bin/nsenter "-t 1 -m -u -n -i -p -- cat /etc/os-release" 127The escape worked; cat failed only because the Talos host has no cat.
The same escape against each selector:
matchNamespaces Pid NotIn host_ns: 0 events (NPOST 0), the hostPID pod is in the host PID namespace
matchNamespaces Mnt NotIn host_ns: 5 events from nsenter + 1 from runc per kubectl exec
Mnt NotIn host_ns and matchBinaries NotIn runc: escape + 3 kubectl exec: 5 events; 3 kubectl exec alone: 04. Secrets are masked before export:
process team-a/debug /bin/sh "-c \"echo run --password=***** --token ***** >/dev/null; sleep 1\""5. Host events reach the export. Host events have no pod field:
kubectl logs -n kube-system tetragon-xxxxx -c export-stdout --tail 400 | grep -v -c '"pod":'The same period on the same node, with each deny list:
chart default deny list: 13 events, 3 without a pod (all /usr/bin/runc)
deny list on this page: 57 events, 33 without a pod (containerd-shim-runc-v2, /pause, udevadm, mdadm, runc, <kernel>)The policy also passes schema validation against the Tetragon v1.7.1 CRDs
(secure-tests/tetragon-runtime-detection-talos/).
Mistakes people make
Keeping the default export deny list
It drops host and kube-system events from the export. Those are the
namespaces with the most privilege. Filter noise with exportAllowList and
field filters instead of dropping whole namespaces.
Selecting on the PID namespace
matchNamespaces with Pid NotIn host_ns looks like "not the host", but a
hostPID pod shares the host's PID namespace. The escape that matters most
then produces no event at all. Use the mount namespace.
Exposing gRPC on TCP
The chart keeps gRPC on a Unix socket by default. Do not change
tetragon.grpc.address to a TCP address without mTLS; anyone who reaches it
can read every event.
Events that never leave the node
The export file is rotated at exportFileMaxSizeMB (10 MB by default). If no
agent ships it, the evidence rolls away within hours on a busy node.
No alert when Tetragon stops
An attacker with root on a node can try to stop the agent. Alert when the DaemonSet is not fully ready or when event counters stop rising.
Secrets in process arguments
Tetragon records full command lines. Add redactionFilters for the flags
your tools use, and fix the tools that take secrets on the command line.
Checklist
- Tetragon runs on every Talos node, 2/2 ready, with the Talos values from its docs.
enableProcessCredandenableProcessNsare on.- The export deny list no longer drops host (
"") andkube-systemevents. - gRPC listens only on the Unix socket.
redactionFiltersmask secret-bearing arguments.- The JSON export is shipped off the node and retained like other security logs.
- A setns policy is applied that selects on the mount namespace, not the PID namespace, and leaves out runc.
- An alert fires when a Tetragon pod is not ready or goes silent.
On an immutable node, the kernel is the only witness that was in the room. Make sure it is taking notes, and that the notes leave the building.
H2-CTDE
Learn it on a live range
Kernel-level detection with eBPF, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on runtime detection and observability
Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.
All runtime detection and observability guides