Runtime detection and observability

Tetragon: catching privilege escalation, the secure way

The container runs as user 1000, the admission policy said so, and the dashboard is green. Then someone finds a setuid binary the base image forgot about, and user 1000 has a very good afternoon.

The short answer

Enable process credentials in Tetragon, then apply a TracingPolicy that hooks the setuid family of system calls with a match on uid 0, and commit_creds rate-limited, both limited to non-host PID namespaces. Alert on every event from a workload that should never change identity. Where every container runs as non-root, add a namespaced Override plus Sigkill policy.

Updated Houssam Hammoudi, CTOTested with Tetragon 1.7.1, Kubernetes 1.34 (kind), kernel 6.8

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Admission controls check a pod's settings when it starts: runAsNonRoot, no privileged mode, capabilities dropped. They do not watch what happens afterward. A process inside a running container can still gain privileges through a setuid binary, a kernel bug, a writable sudo rule baked into the image, or a capability the pod kept.

Privilege escalation is also the step between "attacker runs code in an app" and "attacker controls the node". It is short, and it rarely repeats. If you do not record it when it happens, you do not get a second chance.

Most setups log process starts and nothing else. The moment the process becomes root is not in the logs.

What the docs say

The commit_creds() is a catch all:

Source: Tetragon examples, process-creds-installed.yaml

Sigkill action terminates synchronously the process that made the call that matches the appropriate selectors from the kernel.

Source: Tetragon docs, Selectors, Sigkill action

Override uses the kernel error injection framework and is only available on kernels compiled with CONFIG_BPF_KPROBE_OVERRIDE configuration option.

Source: Tetragon docs, Selectors, Override action

By default, the rate limiting is applied per thread, meaning that only repeated actions by the same thread will be rate limited.

Source: Tetragon docs, Selectors, Rate limiting

The same example file warns that commit_creds fires on every execve, even when nothing changed. On its own it is a firehose. Rate-limit it, limit it to containers, and use the setuid hooks with a match on uid 0 for the events that should page someone.

The secure configuration

First, add credentials to every process event:

bash
helm upgrade tetragon cilium/tetragon -n kube-system --reuse-values \
  --set tetragon.enableProcessCred=true --set tetragon.enableProcessNs=true

Detection across the cluster:

yaml
# Privilege changes inside containers: setuid-family calls and capability changes.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
  name: privilege-changes
spec:
  kprobes:
    # A process asks to become uid 0 (su, sudo, a setuid exploit's payload).
    - call: "sys_setuid"
      syscall: true
      args:
        - index: 0
          type: "int"
      selectors:
        - matchNamespaces:
            - namespace: Pid
              operator: NotIn
              values: ["host_ns"]
          matchArgs:
            - index: 0
              operator: "Equal"
              values: ["0"]
    - call: "sys_setresuid"
      syscall: true
      args:
        - index: 0
          type: "int"
        - index: 1
          type: "int"
        - index: 2
          type: "int"
      selectors:
        - matchNamespaces:
            - namespace: Pid
              operator: NotIn
              values: ["host_ns"]
    # The kernel installs new credentials (covers setuid binaries, capset, user namespaces).
    - call: "commit_creds"
      syscall: false
      args:
        - index: 0
          type: "cred"
      selectors:
        - matchNamespaces:
            - namespace: Pid
              operator: NotIn
              values: ["host_ns"]
          matchActions:
            - action: Post
              rateLimit: "1m"

Enforcement for one namespace whose workloads never need to become root:

yaml
# Enforcement for one namespace: kill any process in team-a pods that asks for uid 0.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicyNamespaced
metadata:
  name: block-setuid-root
  namespace: team-a
spec:
  kprobes:
    - call: "sys_setuid"
      syscall: true
      args:
        - index: 0
          type: "int"
      selectors:
        - matchArgs:
            - index: 0
              operator: "Equal"
              values: ["0"]
          matchActions:
            # Make the call fail, so the kernel never installs root credentials...
            - action: Override
              argError: -1
            # ...and kill the process that tried.
            - action: Sigkill

Override needs a kernel built with CONFIG_BPF_KPROBE_OVERRIDE. Check with grep CONFIG_BPF_KPROBE_OVERRIDE /boot/config-$(uname -r) on a node. Without it, keep Sigkill alone and read the note under "Prove it".

Apply this only to namespaces where every container runs as a non-root user. The reason is in "Mistakes people make".

Turn the events into alerts in your log pipeline, for example:

  • any sys_setuid with argument 0 from a pod whose image runs as non-root;
  • any commit_creds event where euid becomes 0 in a namespace other than your system namespaces;
  • any process exec with CAP_SYS_ADMIN in its effective capabilities outside an allow list of workloads.

Prove it

The lab pods run as uid 1000. An init container planted a setuid-root copy of python3 in a shared volume, the stand-in for a setuid binary the image forgot about:

text
$ kubectl -n team-a logs app -c plant-setuid
-rwsr-xr-x    1 root     root         14000 Sep 17 21:52 /tools/python3
$ kubectl -n team-a exec app -- id
uid=1000 gid=1000 groups=1000

The escalation, once in team-b (detection only) and once in team-a (enforcement, first with Sigkill alone):

text
$ P='import os; os.setuid(0); os.execv("/usr/bin/id", ["id"])'
$ kubectl -n team-b exec app -c app -- /tools/python3 -c "$P"
uid=0(root) gid=1000 groups=1000
$ kubectl -n team-a exec app -c app -- /tools/python3 -c "$P"
command terminated with exit code 137

What Tetragon recorded (tetra getevents -o compact --namespace team-b --namespace team-a):

text
process team-b/app /tools/python3 "-c \"import os; os.setuid(0); os.execv(\"/usr/bin/id\", [\"id\"])\""
setuid  team-b/app /tools/python3 0
syscall team-b/app /tools/python3 commit_creds
process team-b/app /usr/bin/id
exit    team-b/app /usr/bin/id  0
process team-a/app /tools/python3 "-c \"import os; os.setuid(0); os.execv(\"/usr/bin/id\", [\"id\"])\""
setuid  team-a/app /tools/python3 0
setuid  team-a/app /tools/python3 0
syscall team-a/app /tools/python3 commit_creds
exit    team-a/app /tools/python3 "-c \"import os; os.setuid(0); os.execv(\"/usr/bin/id\", [\"id\"])\"" SIGKILL

The same kill in the JSON export, trimmed to the fields that matter:

json
{"function_name":"__x64_sys_setuid","action":"KPROBE_ACTION_SIGKILL","policy_name":"block-setuid-root","args":[{"int_arg":0}],"ns":"team-a"}

Look at the team-a lines again. With Sigkill alone, commit_creds still fires: the kernel finishes the setuid call and installs uid 0, and the process dies when it returns to user space. id never runs, so nothing used the new identity, but the credentials did change. With Override added, the same run shows no commit_creds event: the call fails and the process is killed before anything changes.

Override on its own, to see what the application sees:

text
$ kubectl -n team-a exec app -c app -- /tools/python3 -c "$P"
PermissionError: [Errno 1] Operation not permitted
command terminated with exit code 1

Mistakes people make

Enforcing in a namespace that has root containers

This one looks like success. With the Sigkill policy in team-a and a container that runs as root, every command is killed, even ones that never ask for root:

text
$ kubectl -n team-a exec app -- true
command terminated with exit code 137

The container runtime does it, not your command. When runc starts a process in a container, it switches to the container's user with this call:

go
if err := unix.Setuid(config.UID); err != nil {

Source: runc, libcontainer/init_linux.go, setupUser

For a root container, config.UID is 0, so the policy kills runc's helper before your command starts. A test like kubectl exec app -- su root then "passes" for the wrong reason. Exec probes die the same way: in the lab, a root pod with an exec liveness probe was marked Unhealthy and restarted within 45 seconds. Enforce only where every container sets a non-root runAsUser, and test with a non-root pod.

Watching only process execution

exec events show that su or an exploit ran. They do not show that it worked. Hook the credential change itself.

commit_creds without a rate limit

It fires on every exec in every container. Without rateLimit and a namespace selector it floods the pipeline and gets turned off.

Enforcing cluster-wide on day one

A Sigkill rule for every setuid(0) breaks package managers, init scripts and some web servers that start as root and drop privileges. Detect first, learn what is normal, then enforce per namespace.

Ignoring the host namespace entirely

The examples exclude host_ns to keep noise down. A container escape ends in the host namespace. Keep a separate, narrower policy for host processes, for example setuid calls from binaries outside /usr/bin and /usr/sbin.

Checklist

  • tetragon.enableProcessCred and tetragon.enableProcessNs are on.
  • A cluster-wide policy hooks the setuid family with a match on uid 0.
  • commit_creds is hooked with a namespace selector and a rate limit.
  • Events are shipped off the node and turned into alerts with an allow list.
  • Namespaces that never need root have a namespaced Override plus Sigkill policy.
  • Every container in an enforced namespace runs as a non-root user (runc's own setuid(0) is killed otherwise).
  • Host-namespace privilege changes have their own, narrower policy.
  • Images are scanned for setuid binaries so there is less to escalate with.

Becoming root takes one system call. Make sure that call is the loudest one in the cluster.

H2-CTDE

Learn it on a live range

Kernel-level detection with eBPF, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on runtime detection and observability

Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.

All runtime detection and observability guides