Runtime detection and observability
Alerting when a sensor goes quiet, the secure way
The quietest night your detection stack ever had was the night the attacker turned the sensor off. No alerts, no noise, a perfect dashboard. Silence is a signal too, if you ask for it.
The short answer
Write four rules for every sensor: up == 0 for a failing target, absent_over_time for a target that disappeared, an event rate of zero for an agent that is up but sends nothing, and a comparison with the fleet median for a host that went much quieter than its peers. Unit-test them with promtool test rules.
On this page
What goes wrong
Detection rules watch the data that arrives. When data stops arriving, every rule goes quiet at once, and quiet looks exactly like "nothing bad is happening".
Sensors stop for boring reasons: an agent crashes, a node is replaced and the new one lacks the agent, a relabel change drops the target, a buffer fills up. They also stop for a very interesting reason: an attacker with root on a node kills or unloads the sensor first.
The usual health check, up == 0, only covers one of these cases. It fires
when Prometheus can still see the target and the scrape fails. It says
nothing when the target disappears from service discovery, or when the agent
answers scrapes happily while sending no events.
What the docs say
1 if the instance is healthy, i.e. reachable, or 0 if the scrape failed.
Source: Prometheus docs, Jobs and instances (the up metric)
This is useful for alerting on when no time series exist for a given metric name and label combination for a certain amount of time.
Source: Prometheus docs, Query functions, absent_over_time()
The total number of Tetragon events
Source: Tetragon docs, Metrics reference, tetragon_events_total
up answers "can I scrape it?", not "is it doing its job?". The Tetragon
event counter answers the second question. Other sensors have an equivalent
counter; find it before you need it.
The secure configuration
Four rules, one per failure mode. The examples use Tetragon; replace the job and metric names for Falco, Alloy or any other agent.
# sensor-silence.rules.yml: alert when a security sensor stops reporting.
groups:
- name: sensor-silence
rules:
# 1. The scrape target is gone or failing (agent down, network blocked).
- alert: SensorDown
expr: up{job=~"tetragon|falco|alloy|node"} == 0
for: 5m
labels:
severity: critical
annotations:
summary: "{{ $labels.job }} on {{ $labels.instance }} is not answering scrapes"
runbook: "https://wiki.example.com/runbooks/sensor-silence"
# 2. The series vanished entirely (target removed from discovery, relabel mistake).
# up == 0 cannot fire for a target that no longer exists; absent_over_time can.
- alert: SensorMissing
expr: absent_over_time(up{job="tetragon"}[10m])
labels:
severity: critical
annotations:
summary: "No tetragon target has reported for 10 minutes"
# 3. The agent is up but has stopped producing events (stuck pipeline, full buffer).
- alert: SensorSilent
expr: sum by (instance) (rate(tetragon_events_total[15m])) == 0
for: 15m
labels:
severity: critical
annotations:
summary: "{{ $labels.instance }} is up but has sent no security events for 30 minutes"
# 4. One host went quiet while its peers did not: compare against the fleet.
- alert: SensorMuchQuieterThanPeers
expr: |
sum by (instance) (rate(tetragon_events_total[30m]))
< 0.1 * scalar(quantile(0.5, sum by (instance) (rate(tetragon_events_total[30m]))))
for: 30m
labels:
severity: warning
annotations:
summary: "{{ $labels.instance }} reports less than 10% of the fleet median event rate"Unit tests, one per failure mode, so a later edit cannot break a rule without failing CI:
# promtool unit tests for sensor-silence.rules.yml
rule_files:
- sensor-silence.rules.yml
evaluation_interval: 1m
tests:
- name: a node agent stops answering
interval: 1m
input_series:
- series: 'up{job="tetragon", instance="node-1"}'
values: '1 1 1 0 0 0 0 0 0 0'
alert_rule_test:
- eval_time: 5m
alertname: SensorDown
exp_alerts: [] # down for 2 minutes: not yet (for: 5m)
- eval_time: 9m
alertname: SensorDown
exp_alerts:
- exp_labels: {severity: critical, job: tetragon, instance: node-1}
exp_annotations:
summary: "tetragon on node-1 is not answering scrapes"
runbook: "https://wiki.example.com/runbooks/sensor-silence"
- name: the target disappears from discovery
interval: 1m
input_series:
- series: 'up{job="tetragon", instance="node-1"}'
values: '1 1 1 _ _ _ _ _ _ _ _ _ _ _ _' # _ = no sample at all
alert_rule_test:
- eval_time: 5m
alertname: SensorDown
exp_alerts: [] # nothing to compare: up == 0 never fires
- eval_time: 14m
alertname: SensorMissing
exp_alerts:
- exp_labels: {severity: critical, job: tetragon} # absent_over_time keeps the job="..." matcher as a label
exp_annotations:
summary: "No tetragon target has reported for 10 minutes"
- name: the agent is up but events stop
interval: 1m
input_series:
- series: 'up{job="tetragon", instance="node-1"}'
values: '1x60'
- series: 'tetragon_events_total{instance="node-1", type="PROCESS_EXEC"}'
values: '0+100x10 1000x50' # counter rises for 10 minutes, then flat
alert_rule_test:
- eval_time: 20m
alertname: SensorSilent
exp_alerts: []
- eval_time: 45m
alertname: SensorSilent
exp_alerts:
- exp_labels: {severity: critical, instance: node-1}
exp_annotations:
summary: "node-1 is up but has sent no security events for 30 minutes"
- eval_time: 45m
alertname: SensorDown
exp_alerts: [] # the proof that up alone would miss it
- name: one host far quieter than its peers
interval: 1m
input_series:
- series: 'tetragon_events_total{instance="node-1"}'
values: '0+100x90'
- series: 'tetragon_events_total{instance="node-2"}'
values: '0+100x90'
- series: 'tetragon_events_total{instance="node-3"}'
values: '0+2x90' # 2 events a minute instead of 100
alert_rule_test:
- eval_time: 80m
alertname: SensorMuchQuieterThanPeers
exp_alerts:
- exp_labels: {severity: warning, instance: node-3}
exp_annotations:
summary: "node-3 reports less than 10% of the fleet median event rate"Two more signals from the Tetragon metrics reference are worth an alert of
their own: tetragon_bpf_missed_events_total and
tetragon_observer_ringbuf_events_lost_total rising means the sensor runs
but is dropping events.
The alerting path itself also needs a heartbeat: an always-firing rule
(vector(1)) routed to an external dead man's switch, which pages when the
heartbeat stops.
Prove it
promtool check rules sensor-silence.rules.ymlChecking sensor-silence.rules.yml
SUCCESS: 4 rules foundpromtool test rules sensor-silence.test.yml SUCCESSThe same tests with SensorMissing and SensorSilent removed, leaving only
up-style checks. The target that vanished and the agent that went silent
both pass unnoticed:
FAILED:
name: the target disappears from discovery,
alertname: SensorMissing, time: 14m,
exp:[
0:
Labels:{alertname="SensorMissing", job="tetragon", severity="critical"}
Annotations:{summary="No tetragon target has reported for 10 minutes"}
],
got:[]
name: the agent is up but events stop,
alertname: SensorSilent, time: 45m,
exp:[The test script is secure-tests/alerting-sensor-goes-quiet/run.sh.
Mistakes people make
Only alerting on up == 0
A target removed from discovery has no up series at all, so the comparison
has nothing to compare. Pair it with absent_over_time.
absent() over a whole job
absent_over_time(up{job="tetragon"}[10m]) fires only when every target of
the job is gone. One missing node out of ten does not trigger it. Use the
per-host rate rule or compare the target count with the node count.
Treating the sensor's own metrics as enough
If the attacker stops the metrics endpoint and the event export together,
both go quiet at once. The SensorMissing rule catches that, but only if it
is evaluated somewhere the attacker cannot reach.
Silence alerts with the wrong severity
A sensor going quiet on a production node is a security event, not a maintenance ticket. Route it like one.
Untested rules
A typo in a label name makes a silence rule silent itself. promtool test rules in CI turns that into a failed build.
Checklist
- Every sensor has a
up == 0rule with a shortfor:. - Every sensor job has an
absent_over_timerule. - Every sensor has an "up but no events" rule based on its event counter.
- Hosts are compared with their peers for large drops in event rate.
- Dropped-event counters of the sensor have their own alert.
- The alerting pipeline has a heartbeat to an external dead man's switch.
promtool test rulesruns in CI with a test per rule.- Silence alerts are routed with security severity.
An attacker who cannot hide from your sensors will try to switch them off. Make sure the off switch is the loudest thing in the building.
H2-CTDE
Learn it on a live range
Alerting and incident response, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on runtime detection and observability
Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.
All runtime detection and observability guides