Edge and WAF

Rate limiting login endpoints on Envoy Gateway, the secure way

You rate limited POST /login to five a minute and watched the sixth request get a 429. Then someone sent POST /login?next=/ and got in line again, forever. The limit worked. It was just limiting a different URL.

The short answer

Give the login endpoint its own HTTPRoute (PathPrefix /login, method POST) and attach a BackendTrafficPolicy to that route: a Global rule with a Distinct sourceCIDR for a per-client limit, and a Local rule as a per-replica backstop. Do not rely on a path selector in the rule: it is matched against the path including the query string.

Updated Houssam Hammoudi, CTOTested with Envoy Gateway v1.9.1 (rate limit service with Redis 7.4), Envoy 1.39.1, Kubernetes 1.34 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A login endpoint is where passwords get guessed. A rate limit per client makes guessing slow enough to be useless, and it is one of the cheapest controls you can add at the edge.

Envoy Gateway offers two kinds of limit. A Local limit counts inside each Envoy replica. A Global limit counts in a shared rate limit service backed by Redis. Only the global one can count per client IP across replicas.

Most examples put the limit on the route that serves the whole site and pick out the login with a path selector. Envoy Gateway renders that selector as an exact match on the :path header, and :path includes the query string. POST /login?next=/ and POST /login/ do not match /login, so they are never counted. Most web frameworks route all three to the same login handler.

What the docs say

Local rate limiting does not support distinct matching. If you want to rate limit based on distinct values, you should use Global Rate Limiting.

Source: Envoy Gateway docs, Local Rate Limit

This means that if the data plane has 2 replicas of Envoy running, and the rate limit is 10 requests/second, each replica will allow 10 requests/second.

Source: Envoy Gateway docs, Local Rate Limit

If no client selectors are specified, the rule applies to all traffic of the targeted Route.

Source: Envoy Gateway API reference, RateLimitRule

The API reference describes path as "the request path to match" and does not mention that the query string is part of what is matched.

The secure configuration

A route that carries only login attempts:

yaml
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: login
  namespace: edge
spec:
  parentRefs:
  - name: edge
  hostnames: ["www.example.com"]
  rules:
  - matches:
    - path: {type: PathPrefix, value: /login}   # /login, /login/, /login?x; not /login-help
      method: POST
    backendRefs:
    - name: app
      port: 8080

The limit, attached to that route only:

yaml
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: BackendTrafficPolicy
metadata:
  name: login-rate-limit
  namespace: edge
spec:
  targetRefs:
  - group: gateway.networking.k8s.io
    kind: HTTPRoute
    name: login
  rateLimit:
    # Per client address, shared by all replicas (needs the rate limit
    # service and Redis enabled in the EnvoyGateway configuration).
    global:
      rules:
      - clientSelectors:
        - sourceCIDR: {type: Distinct, value: 0.0.0.0/0}
        limit: {requests: 5, unit: Minute}
      - clientSelectors:
        - sourceCIDR: {type: Distinct, value: "::/0"}
        limit: {requests: 5, unit: Minute}
    # Backstop per replica, for everyone together, if Redis is down or a
    # botnet spreads over many addresses. Multiply by the replica count.
    local:
      rules:
      - limit: {requests: 60, unit: Minute}

sourceCIDR uses the client address Envoy Gateway derives. Set client IP detection first, or every request counts as the load balancer; see IP blocklists.

A per-IP limit does not stop a slow attack spread across many addresses. Pair it with a per-account lockout or delay in the application.

Prove it

What Envoy Gateway renders (offline, no cluster)

The test in secure-tests/rate-limiting-login-endpoints-envoy-gateway/run.sh renders two designs with egctl. Design A is the common example: a rule with method and path selectors on the site-wide route. Design B is the one above, reduced to one local rule so it can run without Redis. The full policy above is checked too.

text
== Envoy Gateway v1.9.1, rendered offline by egctl
-- design A
  BackendTrafficPolicy Accepted="True": Policy has been accepted.
  - name: :method
  exact: POST
  - name: :path
  exact: /login
-- design B
  BackendTrafficPolicy Accepted="True": Policy has been accepted.
  exact: POST
  pathSeparatedPrefix: /login
  fillInterval: 60s
  maxTokens: 5
-- the full policy from the page (global per-IP + local backstop)
  BackendTrafficPolicy Accepted="True": Policy has been accepted.

Design A matches the :path header exactly. Design B matches in the route, where Envoy compares the path without the query string.

The rendered configs on Envoy 1.39.1

The script runs each rendered config on a standalone Envoy and sends six POST /login, then three POST /login?next=/, then two POST /login/:

text
== Design A, one replica: 6 x POST /login, then 3 x POST /login?next=/, then 2 x POST /login/
200 200 200 200 200 429 200 200 200 200 200

== Design B, one replica: 6 x POST /login, then 3 x POST /login?next=/, then 2 x POST /login/
200 200 200 200 200 429 429 429 429 429 429

Design A limits the sixth request and then lets every variant through. Design B keeps limiting.

A local limit is per replica. The same design B on two replicas, with the client going round-robin:

text
== Design B, two replicas behind round-robin: 12 x POST /login
200 200 200 200 200 200 200 200 200 200 429 429

Five per minute became ten. That is why the per-client limit must be global.

On a cluster

The route and policy above on a lab cluster, with backendRefs pointing at a test nginx, the Envoy Gateway rate limit service backed by Redis, and the client IP detection from the blocklist page. Requests went through a proxy that appends X-Forwarded-For, as a cloud load balancer does.

yaml
# Envoy Gateway Helm values: the global limit needs the rate limit service.
config:
  envoyGateway:
    rateLimit:
      backend:
        type: Redis
        redis:
          url: redis.redis-system.svc.cluster.local:6379
text
$ kubectl get backendtrafficpolicy login-rate-limit -n edge -o jsonpath='...'
Accepted=True Policy has been accepted.

One client, one request at a time, reading the Redis counter after each (the lab backend answers 404 to every login):

text
request 1 -> 404   counter=1
request 2 -> 404   counter=2
request 3 -> 404   counter=3
request 4 -> 404   counter=4
request 5 -> 404   counter=5
request 6 -> 429   counter=6
request 7 -> 429   counter=7

The counter key ends in the client's real address (..._remote_address_10.244.2.207_1790304060), so each client has its own budget. A second client in the same minute got its own five. POST /login/ and POST /login?next=/ counted against the same budget. In two runs made right after the rate limit service started, six requests passed before the first 429; every later run gave exactly five.

The window is a calendar minute. The number at the end of the key is the start of the minute. Twelve requests, half a second apart, starting at second 56:

text
02:44:56 404
02:44:57 404
02:44:57 404
02:44:58 404
02:44:58 404
02:44:59 429
02:44:59 429
02:45:00 404
02:45:00 404
02:45:01 404
02:45:01 404
02:45:02 404

Ten attempts in six seconds. A limit of 5 per minute allows up to 10 in any 60 seconds. Size the number with that in mind, and keep the lockout in the application.

If Redis is down, the global limit fails open. With Redis scaled to zero, the rate limit service logged creating redis connection error : dial tcp 10.96.44.34:6379: i/o timeout, and eight POST /login from a fresh client all passed (404 x 8). The local backstop still held: 70 requests in a minute gave 56 x 404 and 14 x 429. That is what the local rule on this page is for; alert on the rate limit service's Redis errors.

On your own cluster:

bash
kubectl get backendtrafficpolicy login-rate-limit -n edge \
  -o jsonpath='{range .status.ancestors[*].conditions[*]}{.type}={.status} {.message}{"\n"}{end}'
for i in $(seq 1 7); do
  curl -s -o /dev/null -w "%{http_code} " -X POST -d "user=test&password=wrong" "https://www.example.com/login?next=/"
done; echo

What you should see: Accepted=True, then five 200 (or whatever your login returns for a wrong password) followed by 429. If you never see 429, the policy is attached to a different route than the one serving /login.

Mistakes people make

Selecting the login path inside a site-wide rule

The path selector compares the whole :path, query string included. Match the login in an HTTPRoute instead and limit the whole route.

Using a Local limit as a per-client limit

Local limits cannot count per client, and each replica has its own bucket. With autoscaling, the effective limit grows with traffic.

Counting the load balancer instead of the client

Without client IP detection, all traffic shares one bucket and one attacker locks out everyone. Configure XFF hops or PROXY protocol first.

Forgetting the other doors

Password reset, magic-link request, MFA verification and token endpoints are guessed the same way. Give each its own route and limit.

Stopping at the edge

A per-IP limit does not see a botnet with one guess per address. The application must still slow down or lock the targeted account.

Checklist

  • Create a dedicated HTTPRoute for POST /login with PathPrefix.
  • Attach the BackendTrafficPolicy to that route, not to the site-wide route.
  • Use a Global rule with sourceCIDR type Distinct for IPv4 and IPv6.
  • Enable the rate limit service with Redis in the EnvoyGateway configuration.
  • Add a Local backstop and size it for the replica count.
  • Configure client IP detection before trusting per-IP counts.
  • Test with ?next=/ and a trailing slash, and expect 429.
  • Repeat for password reset, MFA and token endpoints.

Attackers read your routes more carefully than your dashboards do. Make the limit cover every spelling of the door.

H2-CPQE

Learn it on a live range

WAF, rate limiting and DDoS, in Edge and Post-Quantum Networking: a real host in your browser, and every objective checked on the machine.

Start free

The Dome

Want it run for you?

The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.

See the Dome