DNS and DNSSEC

DNS rate limiting and cookies in Knot, the secure way

Your authoritative server answers everyone, instantly, with more bytes than it was asked for. That is its job, and it is also the exact description of a reflection amplifier someone else aims at a victim.

The short answer

Load mod-cookies and then mod-rrl globally in Knot. RRL limits UDP answers per source address and prefix, and with slip 2 it truncates or drops the excess so real clients retry over TCP. Clients with a valid DNS cookie skip RRL. Share one cookie secret across anycast nodes, and whitelist only your own secondaries.

Updated Houssam Hammoudi, CTOTested with Knot DNS 3.5.4

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

DNS over UDP has no handshake. Anyone can send a query with a forged source address, and the server sends the answer to that address. A signed zone makes the answer much larger than the query. Send many small forged queries to many servers, and a victim receives a flood it never asked for. Your server is the amplifier.

Response rate limiting (RRL) counts answers per source address and prefix. Above the limit, it drops some answers and sends others truncated (TC bit set, no data), so a real client retries over TCP, which cannot be forged. DNS Cookies (RFC 7873) let a real client prove it saw an earlier answer. Knot skips RRL for such clients.

Knot does not load either module by default.

What the docs say

RRL lowers the amplification factor of these attacks by sending some responses as truncated or by dropping them altogether.

Source: Knot DNS 3.5 docs, rrl module

If the Cookies module is active and configured before this module, RRL is not applied to UDP responses with a valid DNS cookie.

Source: Knot DNS 3.5 docs, rrl module

Setting the value to 0 will cause that all rate-limited responses will be dropped.

Source: Knot DNS 3.5 docs, rrl slip

For servers accessed via anycast, to successfully support DNS Cookies, either (1) the server clones must all use the same Server Secret

Source: RFC 7873, Section 6

The Knot docs explain each option in detail. They do not mention the anycast point from RFC 7873: by default each Knot instance generates its own cookie secret, so on an anycast address the cookies stop matching when routing moves a client to another node. The order of the two modules is also easy to get wrong, and nothing warns you.

The secure configuration

yaml
mod-cookies:
  - id: default
    secret: 0x<32 hex characters, the same on every anycast node>
    # To roll the secret: list the new one first and the old one second,
    # on every node, then remove the old one a day later.

mod-rrl:
  - id: default
    rate-limit: 50              # default: per /32 or /128; prefixes get multiples
    instant-limit: 125          # default burst
    slip: 2                     # half truncated, half dropped; never 0
    whitelist: [ 198.51.100.20, 198.51.100.21 ]   # your secondaries and monitors only
    log-period: 30000

template:
  - id: default
    global-module: [ mod-cookies/default, mod-rrl/default ]   # cookies FIRST

Generate the secret once and distribute it with your other secrets:

bash
head -c 16 /dev/urandom | od -An -tx1 | tr -d ' \n' | sed 's/^/0x/'

Watch the counters:

bash
knotc stats mod-rrl        # slipped and dropped
knotc stats mod-cookies    # presence and dropped

Prove it

The test (secure-tests/dns-rate-limiting-cookies-knot/run.sh) uses lab limits of 5 queries per second and an instant limit of 10. Each burst is 100 parallel UDP queries from one loopback address.

Before, with no modules, every query is answered:

text
$ 100 parallel UDP queries from 127.0.0.2 (spoofable source, no cookie)
    100 answered

After, with [ mod-cookies, mod-rrl/default ]:

text
$ 100 parallel UDP queries from 127.0.0.2 (no cookie)
     11 answered
     41 dropped
     48 truncated (TC=1)
bash
knotc stats mod-rrl
text
mod-rrl.slipped = 48
mod-rrl.dropped = 41
notice: module 'mod-rrl/default', address 127.0.0.2 UDP limited on /32 by rate, qname www.example.com.

A real client from the same address still gets an answer over TCP:

bash
kdig @127.0.0.1 -b 127.0.0.2 www.example.com A +tcp +norecurse +short
text
192.0.2.10

A client learns a server cookie. The first reply is BADCOOKIE carrying the server cookie; kdig retries with it and gets NOERROR:

bash
kdig @127.0.0.1 -b 127.0.0.3 www.example.com A +cookie +norecurse
text
;; ->>HEADER<<- opcode: QUERY; status: BADCOOKIE; id: 2547
;; Version: 0; flags: ; UDP size: 1232 B; ext-rcode: BADCOOKIE
;; COOKIE: D9B9AA70A2B0A388010000006AB592A9F1912DDFE84265D7
;; ->>HEADER<<- opcode: QUERY; status: NOERROR; id: 17309
;; COOKIE: D9B9AA70A2B0A388010000006AB592A9F1912DDFE84265D7

With that valid cookie, a burst from 127.0.0.3 is not rate limited:

text
$ 100 parallel UDP queries from 127.0.0.3 with the valid cookie (+cookie=D9B9AA70A2B0A388010000006AB592A9F1912DDFE84265D7)
    100 answered

The same cookie used from another address is rejected, because the server cookie is bound to the client address:

text
$ the same cookie from another address (127.0.0.4): the server cookie does not match
;; WARNING: bad cookie from 127.0.0.1@53(UDP), retrying with the received one
;; ->>HEADER<<- opcode: QUERY; status: BADCOOKIE; id: 25583

The whitelisted secondary is never limited:

text
$ 100 parallel UDP queries from 127.0.0.5 (whitelisted secondary)
    100 answered

Mistakes people make

Loading RRL before cookies

The docs are precise: cookies bypass RRL only if the cookies module comes first. With the order reversed, clients that behave well are limited like forged traffic.

slip: 0

Dropping everything above the limit feels strongest. It removes the TC answers that tell real clients behind a busy resolver to retry over TCP, and the docs say plainly that legitimate requestors "will face denial of service". Keep slip: 2 or the default 1.

Without an explicit secret, each instance makes its own. Behind anycast, a client that moves between nodes keeps getting BADCOOKIE, and RFC 7873 tells it to fall back to TCP. Set one secret on all nodes and roll it with the two-secret form.

Whitelisting large public ranges

Whitelisting a big resolver provider's range removes RRL for exactly the addresses an attacker would forge. Whitelist only addresses you operate.

Tuning limits from a lab test

The lab uses 5 queries per second to make the effect visible. Production limits must fit real resolvers, which send many queries from one address. Start with the defaults and dry-run: on, watch the counters for a week, then enforce.

Checklist

  • Load mod-cookies and then mod-rrl in global-module, in that order.
  • Set slip to 1 or 2, never 0.
  • Set the same 16-byte cookie secret on every node behind an anycast address.
  • Roll the cookie secret with the new and old secrets listed together.
  • Whitelist only your own secondaries and monitoring addresses.
  • Start new limits with dry-run: on and read knotc stats mod-rrl.
  • Graph mod-rrl.slipped, mod-rrl.dropped and mod-cookies.presence.
  • Confirm TCP answers still work from a limited address.

RRL and cookies do not make your server faster or your zone safer. They make sure the traffic your server sends is traffic someone actually asked for, which is the least a good neighbor can do.

H2-CPQE

Learn it on a live range

DNSSEC with Knot, in Edge and Post-Quantum Networking: a real host in your browser, and every objective checked on the machine.

Start free

The Dome

Want it run for you?

The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.

See the Dome