Runtime detection and observability
Security alerting as code in Grafana, the secure way
The alert that would have caught the intrusion was muted in March by someone who found it noisy, from the UI, with no review and no record anyone reads. Detections that anyone can click off are suggestions.
The short answer
Keep security alert rules, contact points and notification policies as provisioning files in git, reviewed like code. Grafana marks file-provisioned rules with provenance file and refuses UI and API edits and deletes. Set execErrState to Error so a broken query alerts, and test each rule against sample data.
On this page
What goes wrong
Alert rules created in the Grafana UI live in Grafana's database. Anyone with the editor role on the folder can change a threshold, add a silence, or delete the rule. There is no review, and the change is easy to miss.
For operations alerts that is an annoyance. For security detections it is a gap an insider or an attacker with a stolen account can use: turn off the rule, then do the thing the rule would have caught.
UI-built rules also drift between environments, cannot be tested before they ship, and disappear with the database if nobody backed it up.
What the docs say
You cannot edit provisioned resources from files in Grafana. You can only change the resource properties by changing the provisioning file and restarting Grafana or carrying out a hot reload.
Source: Grafana docs, Use configuration files to provision alerting resources
Since the policy tree is a single resource, provisioning it will overwrite all policies in the notification policy tree.
Source: Grafana docs, Use configuration files to provision alerting resources
The first quote is the security property this page is about. The second is the trap: provisioning a policy file replaces every routing rule someone made in the UI. Move the whole tree to git at once, or routes will vanish.
The secure configuration
A directory in git, mounted read-only into Grafana at
/etc/grafana/provisioning/:
# provisioning/datasources/loki.yaml
apiVersion: 1
datasources:
- name: Loki (security)
uid: loki-security
type: loki
access: proxy
url: http://loki:3100
editable: false# provisioning/alerting/security-rules.yaml
# Security alert rules, provisioned from git. Grafana will not let anyone edit them in the UI.
apiVersion: 1
groups:
- orgId: 1
name: ssh-sign-in
folder: Security detections
interval: 10s
rules:
- uid: ssh-login-after-failures
title: SSH login after repeated failures
condition: C
data:
# A: failures and successes joined per source (LogQL, instant query over 15m)
- refId: A
datasourceUid: loki-security
relativeTimeRange: {from: 900, to: 0}
model:
refId: A
queryType: instant
expr: >-
(sum by (host, src) (count_over_time({job="sshd"} |= "Failed password" | regexp `from (?P<src>\S+) port` [15m])) > 5)
and on (host, src)
(sum by (host, src) (count_over_time({job="sshd"} |= "Accepted" | regexp `from (?P<src>\S+) port` [15m])) > 0)
# B: one number per series
- refId: B
datasourceUid: __expr__
model: {refId: B, type: reduce, expression: A, reducer: last}
# C: the condition
- refId: C
datasourceUid: __expr__
model: {refId: C, type: threshold, expression: B, conditions: [{evaluator: {type: gt, params: [0]}}]}
# No data means "no attacker got in", which is the normal state.
noDataState: OK
# A broken query must not look like a quiet night.
execErrState: Error
for: 0s
labels:
severity: critical
team: security
annotations:
summary: "{{ $labels.src }} logged in to {{ $labels.host }} after repeated failures"
runbook_url: https://wiki.example.com/runbooks/ssh-login-after-failures# provisioning/alerting/contact-points.yaml
apiVersion: 1
contactPoints:
- orgId: 1
name: security-pager
receivers:
- uid: security-pager-webhook
type: webhook
settings:
url: http://alert-sink:8080/security
httpMethod: POST
disableResolveMessage: false# provisioning/alerting/policies.yaml
apiVersion: 1
policies:
- orgId: 1
receiver: security-pager # default receiver
group_by: [alertname, host]
group_wait: 5s
group_interval: 1m
repeat_interval: 4h
routes:
- receiver: security-pager
object_matchers:
- [team, "=", security]
- [severity, "=", critical]
group_wait: 5sIn a real deployment, the webhook URL or pager key comes from a secret, not
from the file in git: Grafana looks up environment variables in all
provisioning files (url: $SECURITY_PAGER_URL), and the variable is set from
a secret store. Changes go through a pull request, CI runs the rule queries
against sample data, and a restart or a reload through the Admin API applies
them.
Prove it
Grafana loaded the rule from the file and marks it provenance: file:
curl -s -u admin:*** http://grafana:3000/api/v1/provisioning/alert-rules/ssh-login-after-failuresuid: ssh-login-after-failures
title: SSH login after repeated failures
folder rule group: ssh-sign-in
provenance: file
labels: {'severity': 'critical', 'team': 'security'}Even the admin cannot lower the severity through the API:
curl -s -u admin:*** -X PUT -H "Content-Type: application/json" --data-binary @weakened-rule.json \
http://grafana:3000/api/v1/provisioning/alert-rules/ssh-login-after-failures{"statusCode":409,"messageId":"alerting.provenanceMismatch","message":"cannot update with provided provenance 'api', needs 'file'","extra":{"Operation":"update","ProvidedProvenance":"api","StoredProvenance":"file"}}
HTTP 409Or delete it:
curl -s -u admin:*** -X DELETE http://grafana:3000/api/v1/provisioning/alert-rules/ssh-login-after-failures{"statusCode":409,"messageId":"alerting.provenanceMismatch","message":"cannot delete with provided provenance 'classic-api-provisioning', needs 'classic-file-provisioning'","extra":{"Operation":"delete","ProvidedProvenance":"classic-api-provisioning","StoredProvenance":"classic-file-provisioning"}}
HTTP 409End to end: eight failed logins and then a success from 192.0.2.44 were
pushed into Loki. The rule fired, and the webhook received it through the
provisioned policy and contact point:
rule state: SSH login after repeated failures -> firing | health: ok
/security firing SSH login after repeated failures src=192.0.2.44 team=security | 192.0.2.44 logged in to web-1 after repeated failuresDuring the first test run, the grafana/grafana:13.2.2-slim image did not
include the Loki data source plugin. The rule could not run, and because
execErrState: Error was set, Grafana sent a DatasourceError alert to the
same webhook instead of staying quiet. That is the behavior you want. The
test script is secure-tests/security-alerting-as-code-grafana/run.sh.
Mistakes people make
Security rules in the UI
They can be edited or deleted by anyone with folder edit rights, without review. File provisioning makes git the only way to change them.
execErrState and noDataState left at defaults without thought
A query that fails should raise an error alert, not look like a quiet night. "No data" can be normal (no attacker) or a problem (no logs); decide per rule, and cover missing data with a separate sensor-silence rule.
Provisioning part of the policy tree
The policy file replaces the whole tree. Export the existing tree first (Alerting, Notification policies, Export), commit it, then manage it only in git.
Secrets in the provisioning files
Webhook URLs with tokens and pager keys in git are leaked credentials. Reference files or environment variables instead.
No test data
A rule that never fires looks exactly like a rule that works. Keep sample log lines for each detection and check in CI that the query matches them.
Checklist
- Security alert rules exist only as provisioning files in git.
GET /api/v1/provisioning/alert-rules/<uid>showsprovenance: filefor each one.- Contact points and the full notification policy tree are provisioned too.
- Every rule sets
execErrState: Errorand a deliberatenoDataState. - Every rule has a runbook link and a
teamandseveritylabel. - Secrets in contact points come from files or environment variables.
- Changes are reviewed in pull requests, and CI checks each rule against sample data.
- The provisioning directory is mounted read-only.
A detection that anyone can switch off in two clicks protects you from everyone except the people you most need it for. Put it in git.
H2-CTDE
Learn it on a live range
Alerting and incident response, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on runtime detection and observability
Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.
All runtime detection and observability guides