Runtime detection and observability

Alerting when a sensor goes quiet, the secure way

The quietest night your detection stack ever had was the night the attacker turned the sensor off. No alerts, no noise, a perfect dashboard. Silence is a signal too, if you ask for it.

The short answer

Write four rules for every sensor: up == 0 for a failing target, absent_over_time for a target that disappeared, an event rate of zero for an agent that is up but sends nothing, and a comparison with the fleet median for a host that went much quieter than its peers. Unit-test them with promtool test rules.

Updated Houssam Hammoudi, CTOTested with Prometheus promtool 3.14.0

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Detection rules watch the data that arrives. When data stops arriving, every rule goes quiet at once, and quiet looks exactly like "nothing bad is happening".

Sensors stop for boring reasons: an agent crashes, a node is replaced and the new one lacks the agent, a relabel change drops the target, a buffer fills up. They also stop for a very interesting reason: an attacker with root on a node kills or unloads the sensor first.

The usual health check, up == 0, only covers one of these cases. It fires when Prometheus can still see the target and the scrape fails. It says nothing when the target disappears from service discovery, or when the agent answers scrapes happily while sending no events.

What the docs say

1 if the instance is healthy, i.e. reachable, or 0 if the scrape failed.

Source: Prometheus docs, Jobs and instances (the up metric)

This is useful for alerting on when no time series exist for a given metric name and label combination for a certain amount of time.

Source: Prometheus docs, Query functions, absent_over_time()

The total number of Tetragon events

Source: Tetragon docs, Metrics reference, tetragon_events_total

up answers "can I scrape it?", not "is it doing its job?". The Tetragon event counter answers the second question. Other sensors have an equivalent counter; find it before you need it.

The secure configuration

Four rules, one per failure mode. The examples use Tetragon; replace the job and metric names for Falco, Alloy or any other agent.

yaml
# sensor-silence.rules.yml: alert when a security sensor stops reporting.
groups:
  - name: sensor-silence
    rules:
      # 1. The scrape target is gone or failing (agent down, network blocked).
      - alert: SensorDown
        expr: up{job=~"tetragon|falco|alloy|node"} == 0
        for: 5m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.job }} on {{ $labels.instance }} is not answering scrapes"
          runbook: "https://wiki.example.com/runbooks/sensor-silence"

      # 2. The series vanished entirely (target removed from discovery, relabel mistake).
      #    up == 0 cannot fire for a target that no longer exists; absent_over_time can.
      - alert: SensorMissing
        expr: absent_over_time(up{job="tetragon"}[10m])
        labels:
          severity: critical
        annotations:
          summary: "No tetragon target has reported for 10 minutes"

      # 3. The agent is up but has stopped producing events (stuck pipeline, full buffer).
      - alert: SensorSilent
        expr: sum by (instance) (rate(tetragon_events_total[15m])) == 0
        for: 15m
        labels:
          severity: critical
        annotations:
          summary: "{{ $labels.instance }} is up but has sent no security events for 30 minutes"

      # 4. One host went quiet while its peers did not: compare against the fleet.
      - alert: SensorMuchQuieterThanPeers
        expr: |
          sum by (instance) (rate(tetragon_events_total[30m]))
            < 0.1 * scalar(quantile(0.5, sum by (instance) (rate(tetragon_events_total[30m]))))
        for: 30m
        labels:
          severity: warning
        annotations:
          summary: "{{ $labels.instance }} reports less than 10% of the fleet median event rate"

Unit tests, one per failure mode, so a later edit cannot break a rule without failing CI:

yaml
# promtool unit tests for sensor-silence.rules.yml
rule_files:
  - sensor-silence.rules.yml
evaluation_interval: 1m
tests:
  - name: a node agent stops answering
    interval: 1m
    input_series:
      - series: 'up{job="tetragon", instance="node-1"}'
        values: '1 1 1 0 0 0 0 0 0 0'
    alert_rule_test:
      - eval_time: 5m
        alertname: SensorDown
        exp_alerts: []                      # down for 2 minutes: not yet (for: 5m)
      - eval_time: 9m
        alertname: SensorDown
        exp_alerts:
          - exp_labels: {severity: critical, job: tetragon, instance: node-1}
            exp_annotations:
              summary: "tetragon on node-1 is not answering scrapes"
              runbook: "https://wiki.example.com/runbooks/sensor-silence"

  - name: the target disappears from discovery
    interval: 1m
    input_series:
      - series: 'up{job="tetragon", instance="node-1"}'
        values: '1 1 1 _ _ _ _ _ _ _ _ _ _ _ _'   # _ = no sample at all
    alert_rule_test:
      - eval_time: 5m
        alertname: SensorDown
        exp_alerts: []                      # nothing to compare: up == 0 never fires
      - eval_time: 14m
        alertname: SensorMissing
        exp_alerts:
          - exp_labels: {severity: critical, job: tetragon}   # absent_over_time keeps the job="..." matcher as a label
            exp_annotations:
              summary: "No tetragon target has reported for 10 minutes"

  - name: the agent is up but events stop
    interval: 1m
    input_series:
      - series: 'up{job="tetragon", instance="node-1"}'
        values: '1x60'
      - series: 'tetragon_events_total{instance="node-1", type="PROCESS_EXEC"}'
        values: '0+100x10 1000x50'          # counter rises for 10 minutes, then flat
    alert_rule_test:
      - eval_time: 20m
        alertname: SensorSilent
        exp_alerts: []
      - eval_time: 45m
        alertname: SensorSilent
        exp_alerts:
          - exp_labels: {severity: critical, instance: node-1}
            exp_annotations:
              summary: "node-1 is up but has sent no security events for 30 minutes"
      - eval_time: 45m
        alertname: SensorDown
        exp_alerts: []                      # the proof that up alone would miss it

  - name: one host far quieter than its peers
    interval: 1m
    input_series:
      - series: 'tetragon_events_total{instance="node-1"}'
        values: '0+100x90'
      - series: 'tetragon_events_total{instance="node-2"}'
        values: '0+100x90'
      - series: 'tetragon_events_total{instance="node-3"}'
        values: '0+2x90'                    # 2 events a minute instead of 100
    alert_rule_test:
      - eval_time: 80m
        alertname: SensorMuchQuieterThanPeers
        exp_alerts:
          - exp_labels: {severity: warning, instance: node-3}
            exp_annotations:
              summary: "node-3 reports less than 10% of the fleet median event rate"

Two more signals from the Tetragon metrics reference are worth an alert of their own: tetragon_bpf_missed_events_total and tetragon_observer_ringbuf_events_lost_total rising means the sensor runs but is dropping events.

The alerting path itself also needs a heartbeat: an always-firing rule (vector(1)) routed to an external dead man's switch, which pages when the heartbeat stops.

Prove it

bash
promtool check rules sensor-silence.rules.yml
text
Checking sensor-silence.rules.yml
  SUCCESS: 4 rules found
bash
promtool test rules sensor-silence.test.yml
text
  SUCCESS

The same tests with SensorMissing and SensorSilent removed, leaving only up-style checks. The target that vanished and the agent that went silent both pass unnoticed:

text
  FAILED:
    name: the target disappears from discovery,
    alertname: SensorMissing, time: 14m, 
        exp:[
            0:
              Labels:{alertname="SensorMissing", job="tetragon", severity="critical"}
              Annotations:{summary="No tetragon target has reported for 10 minutes"}
            ], 
        got:[]
    name: the agent is up but events stop,
    alertname: SensorSilent, time: 45m, 
        exp:[

The test script is secure-tests/alerting-sensor-goes-quiet/run.sh.

Mistakes people make

Only alerting on up == 0

A target removed from discovery has no up series at all, so the comparison has nothing to compare. Pair it with absent_over_time.

absent() over a whole job

absent_over_time(up{job="tetragon"}[10m]) fires only when every target of the job is gone. One missing node out of ten does not trigger it. Use the per-host rate rule or compare the target count with the node count.

Treating the sensor's own metrics as enough

If the attacker stops the metrics endpoint and the event export together, both go quiet at once. The SensorMissing rule catches that, but only if it is evaluated somewhere the attacker cannot reach.

Silence alerts with the wrong severity

A sensor going quiet on a production node is a security event, not a maintenance ticket. Route it like one.

Untested rules

A typo in a label name makes a silence rule silent itself. promtool test rules in CI turns that into a failed build.

Checklist

  • Every sensor has a up == 0 rule with a short for:.
  • Every sensor job has an absent_over_time rule.
  • Every sensor has an "up but no events" rule based on its event counter.
  • Hosts are compared with their peers for large drops in event rate.
  • Dropped-event counters of the sensor have their own alert.
  • The alerting pipeline has a heartbeat to an external dead man's switch.
  • promtool test rules runs in CI with a test per rule.
  • Silence alerts are routed with security severity.

An attacker who cannot hide from your sensors will try to switch them off. Make sure the off switch is the loudest thing in the building.

H2-CTDE

Learn it on a live range

Alerting and incident response, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on runtime detection and observability

Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.

All runtime detection and observability guides