Runtime detection and observability

Tetragon egress monitoring for CI, the secure way

Your build pulls dependencies from three registries. Last Tuesday it also talked to a server in a country none of your vendors are in, for eleven seconds, from a postinstall script. The network policy allowed it because it allowed "the internet".

The short answer

Apply a namespaced Tetragon TracingPolicy in the CI namespace that hooks tcp_connect and reports connections outside your private ranges, with the binary, arguments and pod that made each one. Build an allow list from a week of events, then enforce by failing connect() at the kernel security hook, plus a network policy, for namespaces with a stable allow list.

Updated Houssam Hammoudi, CTOTested with Tetragon 1.7.1, Cilium 1.20.2, Kubernetes 1.34 (kind), kernel 6.8

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

CI jobs run code nobody reviewed line by line: dependency install scripts, build plugins, test fixtures. A compromised package can read the job's secrets and send them anywhere the build can reach.

Network policies can block egress, but many CI namespaces allow the whole internet because builds need "some registries". And flow logs show an IP address and a pod, not which of the hundred processes in a build made the connection.

Tetragon sees the connection at the kernel with the process attached: the binary, its arguments and its parent. That is the difference between "the build pod talked to 203.0.113.50" and "node running postinstall.js from package X talked to 203.0.113.50".

What the docs say

Once you have this information, you can customize a policy to exclude network traffic to the networks stored in the PODCIDR and SERVICECIDR environment variables.

Source: Tetragon docs, Network Monitoring

Sigkill action terminates synchronously the process that made the call that matches the appropriate selectors from the kernel.

Source: Tetragon docs, Selectors, Sigkill action

Starting from kernel version 5.7 overriding security_ hooks is also possible.

Source: Tetragon docs, Selectors, Override action

The docs show the cluster-wide version. For CI, a namespaced policy with a pod selector is better: it watches only build pods, so the event volume stays small and every event is worth reading.

The secure configuration

Detection in the CI namespace, for build pods only:

yaml
# Outbound TCP connections from CI build pods that leave the cluster's private ranges.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicyNamespaced
metadata:
  name: ci-egress
  namespace: ci
spec:
  podSelector:
    matchLabels:
      role: build
  kprobes:
    - call: "tcp_connect"
      syscall: false
      args:
        - index: 0
          type: "sock"
      selectors:
        - matchArgs:
            - index: 0
              operator: "NotDAddr"
              values:
                - "10.0.0.0/8"        # cluster and VPC ranges: adjust to yours
                - "172.16.0.0/12"
                - "192.168.0.0/16"
                - "127.0.0.0/8"

Enforcement for a namespace whose builds only talk to the cluster and one package mirror:

yaml
# The same rule in enforcement mode, for one namespace with a known allow list.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicyNamespaced
metadata:
  name: ci-egress-enforce
  namespace: ci-locked
spec:
  kprobes:
    - call: "tcp_connect"
      syscall: false
      args:
        - index: 0
          type: "sock"
      selectors:
        - matchArgs:
            - index: 0
              operator: "NotDAddr"
              values:
                - "10.0.0.0/8"
                - "127.0.0.0/8"
                - "198.51.100.10/32"  # the package mirror
          matchActions:
            - action: Sigkill

Sigkill on tcp_connect kills the build, but only after the kernel has sent the SYN: the TCP handshake with the outside host completes, then the socket closes (Prove it, step 3). To stop the connection before any packet leaves, fail connect() itself at the kernel's security hook. This follows Tetragon's own security-socket-connect-block-others.yaml example:

yaml
# Block outbound connections before any packet leaves: fail connect() itself.
apiVersion: cilium.io/v1alpha1
kind: TracingPolicyNamespaced
metadata:
  name: ci-egress-block
  namespace: ci-locked
spec:
  options:
    - name: "disable-kprobe-multi"   # Override on a security_ hook is refused through kprobe_multi
      value: "1"
  kprobes:
    - call: "security_socket_connect"
      syscall: false
      args:
        - index: 1
          type: "sockaddr"
      selectors:
        - matchArgs:
            - index: 1
              operator: "Family"
              values: ["AF_INET"]
            - index: 1
              operator: "NotSAddr"
              values:
                - "10.0.0.0/8"
                - "127.0.0.0/8"
          matchActions:
            - action: Override
              argError: -1```

`Override` on a `security_` hook is refused through the fast `kprobe_multi`
attach, so the policy turns it off for its own hooks with the
`disable-kprobe-multi` option. It also needs a kernel with
`CONFIG_BPF_KPROBE_OVERRIDE`. The `Family` match keeps it to IPv4; add an
`AF_INET6` selector with your IPv6 ranges if pods have IPv6.

```bash
kubectl apply -f ci-egress.yaml
# after a week of events and a reviewed allow list:
kubectl apply -f ci-egress-block.yaml
# check that it loaded: kubectl apply succeeds even when Tetragon rejects the policy
kubectl exec -n kube-system ds/tetragon -c tetragon -- tetra tracingpolicy list

Use the events to build the allow list, then put it in a network policy as well. Tetragon shows who connected; the network policy stops the packets at the network layer as a second, independent control.

Prove it

Run on a lab cluster. ci has build-123 (label role=build) and tool (label role=tool); ci-locked has build-9. In-cluster traffic went to a Service in 10.96.0.0/16, inside the excluded 10.0.0.0/8.

1. Detection sees build pods leaving the cluster, and nothing else:

text
$ kubectl -n ci exec build-123 -- sh -c 'curl ... https://example.com; curl ... http://web.edge-test.svc.cluster.local/'
example.com 200
in-cluster 200
$ kubectl -n ci exec tool -- curl ... https://example.com
example.com 200
$ tetra getevents -o compact --namespace ci
process ci/build-123 /usr/bin/curl "-s -m 5 -o /dev/null -w \"example.com %{http_code}\\n\" https://example.com"
connect ci/build-123 /usr/bin/curl tcp 10.244.2.174:60136 -> 104.20.23.154:443
exit    ci/build-123 /usr/bin/curl "... https://example.com" 0
process ci/build-123 /usr/bin/curl "... http://web.edge-test.svc.cluster.local/"
exit    ci/build-123 /usr/bin/curl "... http://web.edge-test.svc.cluster.local/" 0
process ci/tool /usr/bin/curl "... https://example.com"
exit    ci/tool /usr/bin/curl "... https://example.com" 0

One connect event: the build pod to example.com. The in-cluster call and the pod without role=build produce none.

2. Sigkill on tcp_connect kills the build:

text
$ kubectl -n ci-locked exec build-9 -- curl ... https://example.com
command terminated with exit code 137
connect ci-locked/build-9 /usr/bin/curl tcp 10.244.2.103:43878 -> 104.20.23.154:443
exit    ci-locked/build-9 /usr/bin/curl "... https://example.com" SIGKILL

3. But the handshake already happened. Hubble, for the same kind of request:

text
ci-locked/build-9:55630 (ID:29684) -> 104.20.23.154:443 (ID:16777226) to-stack FORWARDED (TCP Flags: SYN)
ci-locked/build-9:55630 (ID:29684) -> 104.20.23.154:443 (ID:16777226) to-stack FORWARDED (TCP Flags: ACK)
ci-locked/build-9:55630 (ID:29684) -> 104.20.23.154:443 (ID:16777226) to-stack FORWARDED (TCP Flags: ACK, FIN)

No TLS and no data left, but the outside host saw a full TCP connection from your cluster. Enough to confirm that code ran, or to signal through which ports or hosts get connected to.

4. ci-egress-block stops it before the SYN:

text
$ kubectl -n ci-locked exec build-9 -- curl -sS -m 5 -o /dev/null https://example.com
curl: (7) Failed to connect to example.com port 443 after 14 ms: Could not connect to server
$ kubectl -n ci-locked exec build-9 -- curl ... http://web.edge-test.svc.cluster.local/
in-cluster 200
$ hubble observe --from-pod ci-locked/build-9 --since 15s | grep ':443'
$ tetra getevents -o compact --namespace ci-locked
process ci-locked/build-9 /usr/bin/curl "-sS -m 5 -o /dev/null -w \"example.com %{http_code}\\n\" https://example.com"
syscall ci-locked/build-9 /usr/bin/curl security_socket_connect
syscall ci-locked/build-9 /usr/bin/curl security_socket_connect
syscall ci-locked/build-9 /usr/bin/curl security_socket_connect
syscall ci-locked/build-9 /usr/bin/curl security_socket_connect
exit    ci-locked/build-9 /usr/bin/curl "... https://example.com" 7

Hubble printed nothing: no packet to port 443 left the pod. Four blocked calls, because curl tried each of example.com's addresses. The build fails with an ordinary connection error instead of a kill.

5. Check that the policy loaded. Without the disable-kprobe-multi option, kubectl apply still said created, the object had no status, and the build kept reaching the internet. Only Tetragon knew:

text
$ tetra tracingpolicy list
ID   NAME                    STATE        FILTERID   NAMESPACE   SENSORS          KERNELMEMORY   MODE           NPOST   NENFORCE   NMONITOR
12   ci-egress-block         load_error   12         ci-locked                    0 B            unknown        0       0          0
$ kubectl -n kube-system logs ds/tetragon -c tetragon | grep ci-egress-block
level=warn msg="adding tracing policy failed" error="policy handler 'tracing' failed loading policy 'ci-egress-block': override validation failed: can't override 'security_socket_connect' function with kprobe_multi, use --disable-kprobe-multi option"

With the option, the same list shows enabled and MODE enforce.

Mistakes people make

Trusting kubectl apply for a TracingPolicy

The API server accepts any policy that fits the CRD schema. Whether Tetragon could load it is a separate question, answered only by tetra tracingpolicy list (STATE enabled) and the agent log. Check after every change, on every node.

Believing Sigkill stops the connection

On tcp_connect, the kill lands after the SYN is sent. Use Override on security_socket_connect to fail the call, and a network policy to drop the packets.

Watching IPs without processes

A flow log with a pod name says the build talked to an address. It does not say which dependency did it. Keep the process fields in the exported events.

Enforcing before learning

Builds reach more places than anyone remembers: mirrors, license servers, telemetry. Run in detection mode long enough to see a full release cycle.

Ranges copied from the example

The Tetragon example uses all of RFC 1918. If your cluster or VPC uses different ranges, or if some private ranges are other tenants, list your own.

Forgetting DNS and UDP

tcp_connect does not see UDP. DNS tunneling and QUIC need their own controls: a DNS proxy with an allow list, and network policy for UDP.

Checklist

  • A namespaced tcp_connect policy runs in every CI namespace, selecting build pods.
  • The excluded ranges match your own cluster and VPC ranges.
  • Events include the binary and arguments, and are shipped off the node.
  • An allow list of destinations is built from observed events and reviewed.
  • Namespaces with a stable allow list have an enforcement policy and a network policy.
  • DNS and UDP egress are controlled separately.
  • New unexpected destinations raise an alert tied to the job that made them.

A build that can talk to anyone will, eventually, talk to someone you would rather it did not. Know who is on the line.

H2-CTDE

Learn it on a live range

Network visibility, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on runtime detection and observability

Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.

All runtime detection and observability guides