Runtime detection and observability
Tetragon: catching privilege escalation, the secure way
The container runs as user 1000, the admission policy said so, and the dashboard is green. Then someone finds a setuid binary the base image forgot about, and user 1000 has a very good afternoon.
The short answer
Enable process credentials in Tetragon, then apply a TracingPolicy that hooks the setuid family of system calls with a match on uid 0, and commit_creds rate-limited, both limited to non-host PID namespaces. Alert on every event from a workload that should never change identity. Where every container runs as non-root, add a namespaced Override plus Sigkill policy.
On this page
What goes wrong
Admission controls check a pod's settings when it starts: runAsNonRoot,
no privileged mode, capabilities dropped. They do not watch what happens
afterward. A process inside a running container can still gain privileges
through a setuid binary, a kernel bug, a writable sudo rule baked into the
image, or a capability the pod kept.
Privilege escalation is also the step between "attacker runs code in an app" and "attacker controls the node". It is short, and it rarely repeats. If you do not record it when it happens, you do not get a second chance.
Most setups log process starts and nothing else. The moment the process becomes root is not in the logs.
What the docs say
The commit_creds() is a catch all:
Source: Tetragon examples, process-creds-installed.yaml
Sigkill action terminates synchronously the process that made the call that matches the appropriate selectors from the kernel.
Source: Tetragon docs, Selectors, Sigkill action
Override uses the kernel error injection framework and is only available on kernels compiled with CONFIG_BPF_KPROBE_OVERRIDE configuration option.
Source: Tetragon docs, Selectors, Override action
By default, the rate limiting is applied per thread, meaning that only repeated actions by the same thread will be rate limited.
Source: Tetragon docs, Selectors, Rate limiting
The same example file warns that commit_creds fires on every execve, even
when nothing changed. On its own it is a firehose. Rate-limit it, limit it to
containers, and use the setuid hooks with a match on uid 0 for the events
that should page someone.
The secure configuration
First, add credentials to every process event:
helm upgrade tetragon cilium/tetragon -n kube-system --reuse-values \
--set tetragon.enableProcessCred=true --set tetragon.enableProcessNs=trueDetection across the cluster:
# Privilege changes inside containers: setuid-family calls and capability changes.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
name: privilege-changes
spec:
kprobes:
# A process asks to become uid 0 (su, sudo, a setuid exploit's payload).
- call: "sys_setuid"
syscall: true
args:
- index: 0
type: "int"
selectors:
- matchNamespaces:
- namespace: Pid
operator: NotIn
values: ["host_ns"]
matchArgs:
- index: 0
operator: "Equal"
values: ["0"]
- call: "sys_setresuid"
syscall: true
args:
- index: 0
type: "int"
- index: 1
type: "int"
- index: 2
type: "int"
selectors:
- matchNamespaces:
- namespace: Pid
operator: NotIn
values: ["host_ns"]
# The kernel installs new credentials (covers setuid binaries, capset, user namespaces).
- call: "commit_creds"
syscall: false
args:
- index: 0
type: "cred"
selectors:
- matchNamespaces:
- namespace: Pid
operator: NotIn
values: ["host_ns"]
matchActions:
- action: Post
rateLimit: "1m"Enforcement for one namespace whose workloads never need to become root:
# Enforcement for one namespace: kill any process in team-a pods that asks for uid 0.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicyNamespaced
metadata:
name: block-setuid-root
namespace: team-a
spec:
kprobes:
- call: "sys_setuid"
syscall: true
args:
- index: 0
type: "int"
selectors:
- matchArgs:
- index: 0
operator: "Equal"
values: ["0"]
matchActions:
# Make the call fail, so the kernel never installs root credentials...
- action: Override
argError: -1
# ...and kill the process that tried.
- action: SigkillOverride needs a kernel built with CONFIG_BPF_KPROBE_OVERRIDE. Check
with grep CONFIG_BPF_KPROBE_OVERRIDE /boot/config-$(uname -r) on a node.
Without it, keep Sigkill alone and read the note under "Prove it".
Apply this only to namespaces where every container runs as a non-root user. The reason is in "Mistakes people make".
Turn the events into alerts in your log pipeline, for example:
- any
sys_setuidwith argument0from a pod whose image runs as non-root; - any
commit_credsevent whereeuidbecomes0in a namespace other than your system namespaces; - any process exec with
CAP_SYS_ADMINin its effective capabilities outside an allow list of workloads.
Prove it
The lab pods run as uid 1000. An init container planted a setuid-root copy
of python3 in a shared volume, the stand-in for a setuid binary the image
forgot about:
$ kubectl -n team-a logs app -c plant-setuid
-rwsr-xr-x 1 root root 14000 Sep 17 21:52 /tools/python3
$ kubectl -n team-a exec app -- id
uid=1000 gid=1000 groups=1000The escalation, once in team-b (detection only) and once in team-a
(enforcement, first with Sigkill alone):
$ P='import os; os.setuid(0); os.execv("/usr/bin/id", ["id"])'
$ kubectl -n team-b exec app -c app -- /tools/python3 -c "$P"
uid=0(root) gid=1000 groups=1000
$ kubectl -n team-a exec app -c app -- /tools/python3 -c "$P"
command terminated with exit code 137What Tetragon recorded (tetra getevents -o compact --namespace team-b --namespace team-a):
process team-b/app /tools/python3 "-c \"import os; os.setuid(0); os.execv(\"/usr/bin/id\", [\"id\"])\""
setuid team-b/app /tools/python3 0
syscall team-b/app /tools/python3 commit_creds
process team-b/app /usr/bin/id
exit team-b/app /usr/bin/id 0
process team-a/app /tools/python3 "-c \"import os; os.setuid(0); os.execv(\"/usr/bin/id\", [\"id\"])\""
setuid team-a/app /tools/python3 0
setuid team-a/app /tools/python3 0
syscall team-a/app /tools/python3 commit_creds
exit team-a/app /tools/python3 "-c \"import os; os.setuid(0); os.execv(\"/usr/bin/id\", [\"id\"])\"" SIGKILLThe same kill in the JSON export, trimmed to the fields that matter:
{"function_name":"__x64_sys_setuid","action":"KPROBE_ACTION_SIGKILL","policy_name":"block-setuid-root","args":[{"int_arg":0}],"ns":"team-a"}Look at the team-a lines again. With Sigkill alone, commit_creds
still fires: the kernel finishes the setuid call and installs uid 0, and
the process dies when it returns to user space. id never runs, so nothing
used the new identity, but the credentials did change. With Override
added, the same run shows no commit_creds event: the call fails and the
process is killed before anything changes.
Override on its own, to see what the application sees:
$ kubectl -n team-a exec app -c app -- /tools/python3 -c "$P"
PermissionError: [Errno 1] Operation not permitted
command terminated with exit code 1Mistakes people make
Enforcing in a namespace that has root containers
This one looks like success. With the Sigkill policy in team-a and a
container that runs as root, every command is killed, even ones that never
ask for root:
$ kubectl -n team-a exec app -- true
command terminated with exit code 137The container runtime does it, not your command. When runc starts a process in a container, it switches to the container's user with this call:
if err := unix.Setuid(config.UID); err != nil {Source: runc, libcontainer/init_linux.go, setupUser
For a root container, config.UID is 0, so the policy kills runc's helper
before your command starts. A test like kubectl exec app -- su root then
"passes" for the wrong reason. Exec probes die the same way: in the lab, a
root pod with an exec liveness probe was marked Unhealthy and restarted
within 45 seconds. Enforce only where every container sets a non-root
runAsUser, and test with a non-root pod.
Watching only process execution
exec events show that su or an exploit ran. They do not show that it
worked. Hook the credential change itself.
commit_creds without a rate limit
It fires on every exec in every container. Without rateLimit and a
namespace selector it floods the pipeline and gets turned off.
Enforcing cluster-wide on day one
A Sigkill rule for every setuid(0) breaks package managers, init scripts
and some web servers that start as root and drop privileges. Detect first,
learn what is normal, then enforce per namespace.
Ignoring the host namespace entirely
The examples exclude host_ns to keep noise down. A container escape ends in
the host namespace. Keep a separate, narrower policy for host processes, for
example setuid calls from binaries outside /usr/bin and /usr/sbin.
Checklist
tetragon.enableProcessCredandtetragon.enableProcessNsare on.- A cluster-wide policy hooks the setuid family with a match on uid 0.
commit_credsis hooked with a namespace selector and a rate limit.- Events are shipped off the node and turned into alerts with an allow list.
- Namespaces that never need root have a namespaced
OverrideplusSigkillpolicy. - Every container in an enforced namespace runs as a non-root user (runc's own setuid(0) is killed otherwise).
- Host-namespace privilege changes have their own, narrower policy.
- Images are scanned for setuid binaries so there is less to escalate with.
Becoming root takes one system call. Make sure that call is the loudest one in the cluster.
H2-CTDE
Learn it on a live range
Kernel-level detection with eBPF, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on runtime detection and observability
Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.
All runtime detection and observability guides