Runtime detection and observability
Detecting sign-in brute force from logs, the secure way
Counting failed logins is easy, and it produces an alert every four minutes forever, because the internet never stops trying. The alert you actually need is the rare one: the attacker who stopped failing.
The short answer
Parse the source address and user name out of sshd logs with LogQL, then alert on three patterns: more than 10 failures from one source in 5 minutes, one source trying five or more user names, and a successful login from a source that failed repeatedly. The last one is critical; route the first two to a dashboard or low-priority channel.
On this page
What goes wrong
Any host with SSH on the internet sees thousands of failed logins a day. An alert on "more than N failures" is technically correct and practically useless: it fires all the time, people mute it, and the channel becomes noise.
Attackers also adapt. Password spraying tries one or two common passwords against many user names, so no single account crosses a lockout threshold. Slow attacks spread attempts over hours.
The event that matters is rare: a source that failed many times, then succeeded. That means a guessed password or a leaked credential. It is often lost among the failures, or never alerted on at all.
What the docs say
The ruler is responsible for continually evaluating a set of configurable queries and performing an action based on the result.
Source: Grafana Loki docs, Alerting and recording rules
We support Prometheus-compatible alerting rules.
Source: Grafana Loki docs, Alerting and recording rules
count_over_time(log-range): counts the entries for each log stream within the given range.
Source: Grafana Loki docs, Metric queries
"For each log stream" is the part to remember. A stream is one set of labels, usually one host. The source address is inside the log line, so the query has to extract it with a parser before it can count per attacker.
The secure configuration
A Loki ruler rule file. The same expressions work as Grafana-managed alert rules against a Loki data source.
# Loki ruler rules for sign-in attacks on SSH. Group per concern; every rule has a runbook.
groups:
- name: ssh-sign-in
rules:
# Many failures from one source: classic brute force.
- alert: SSHBruteForceFromSource
expr: |
sum by (host, src) (
count_over_time({job="sshd"} |= "Failed password" | regexp `from (?P<src>\S+) port` [5m])
) > 10
for: 0m
labels:
severity: warning
annotations:
summary: "{{ $labels.src }} failed SSH password login {{ $value }} times in 5m on {{ $labels.host }}"
runbook: "https://wiki.example.com/runbooks/ssh-brute-force"
# One source trying many different user names: password spraying.
- alert: SSHPasswordSpray
expr: |
count by (host, src) (
sum by (host, src, user) (
count_over_time({job="sshd"} |= "Failed password" | regexp `for (invalid user )?(?P<user>\S+) from (?P<src>\S+) port` [15m])
)
) >= 5
labels:
severity: warning
annotations:
summary: "{{ $labels.src }} tried {{ $value }} different user names on {{ $labels.host }}"
# The one that matters most: a success from a source that was failing.
- alert: SSHLoginAfterFailures
expr: |
(sum by (host, src) (count_over_time({job="sshd"} |= "Failed password" | regexp `from (?P<src>\S+) port` [15m])) > 5)
and on (host, src)
(sum by (host, src) (count_over_time({job="sshd"} |= "Accepted" | regexp `from (?P<src>\S+) port` [15m])) > 0)
labels:
severity: critical
annotations:
summary: "{{ $labels.src }} logged in to {{ $labels.host }} after repeated failures"Make sure the logs have what the rules need:
- sshd logs failures and successes at the default
LogLevel INFO. UseVERBOSEif you also want the key fingerprint of each successful login. - The shipper (Alloy, Promtail, Fluent Bit) must send the auth log or the
journal entries of
sshdwith a stablejobandhostlabel. - Hosts behind a load balancer or proxy must log the real client address, or every line has the proxy's address as the source.
Prove it
The test pushed four patterns: 40 failures from 203.0.113.7, one failure
for each of 8 user names from 198.51.100.23, 8 failures and then a
successful password login from 192.0.2.44, and one normal key login from
192.0.2.10.
Failures per source, extracted from the log lines:
curl -s -G http://loki:3100/loki/api/v1/query --data-urlencode 'query=sum by (src) (count_over_time({job="sshd"} |= "Failed password" | regexp `from (?P<src>\S+) port` [5m]))'src=192.0.2.44 value=8
src=198.51.100.23 value=8
src=203.0.113.7 value=40The three detections, run as instant queries:
== brute force: more than 10 failures from one source
host=web-1 src=203.0.113.7 value=40
== spray: 5 or more user names from one source
host=web-1 src=198.51.100.23 value=8
== success after failures
host=web-1 src=192.0.2.44 value=8192.0.2.44 never crossed the brute-force threshold of 10, and it tried only
one user name. Only the third rule catches it, and it is the only one of the
three sources that got in. The normal key login from 192.0.2.10 triggered
nothing.
The Loki ruler loaded the rule file and all three alerts fired:
curl -s http://loki:3100/prometheus/api/v1/alertsfiring SSHBruteForceFromSource 203.0.113.7 | 203.0.113.7 failed SSH password login 40 times in 5m on web-1
firing SSHLoginAfterFailures 192.0.2.44 | 192.0.2.44 logged in to web-1 after repeated failures
firing SSHPasswordSpray 198.51.100.23 | 198.51.100.23 tried 8 different user names on web-1The test (the query output reduced to labels and values by a small Python
filter) is secure-tests/detect-sign-in-brute-force/run.sh.
Mistakes people make
Paging on every failure count
Brute force against a keys-only SSH server is noise. Page on
SSHLoginAfterFailures; send the others to a dashboard or a daily summary.
Counting per stream instead of per source
count_over_time({job="sshd"} |= "Failed password" [5m]) counts per host.
One attacker and a hundred attackers look the same. Extract the source with
regexp or pattern and aggregate by it.
A window shorter than the attack
A slow attacker making one attempt a minute never crosses "10 in 5 minutes". Add a longer-window rule (for example 50 in 6 hours) for low-and-slow patterns.
Missing the success
Many setups only log or ship failures. Make sure Accepted lines reach Loki
too, or the most important rule has nothing to match.
No test data
A regex that stops matching after a log format change turns every rule silent. Keep a small file of sample lines and run the queries against it in CI, as the test does here.
Checklist
- sshd success and failure lines reach Loki with
jobandhostlabels. - The source address is the real client address, not a proxy.
- A brute-force rule counts failures per source, not per host.
- A spray rule counts distinct user names per source.
- A critical rule fires on a success from a source with recent failures.
- Only the critical rule pages a person.
- A test pushes sample lines and checks that each rule matches.
- Every rule has a runbook link.
Failed logins are the weather report. A login after failures is the knock on your door. Build the alert for the knock.
H2-CTDE
Learn it on a live range
Detection engineering, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on runtime detection and observability
Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.
All runtime detection and observability guides