Service mesh

Retries and non-idempotent requests in a mesh, the secure way

A customer was charged three times, the payment service logs show three requests, and the frontend swears it sent one. Nobody wrote a retry loop. The mesh did, politely, because the first answer was a 504.

The short answer

A mesh proxy retries requests the application never asked to repeat. Kuma's default mesh retries HTTP up to five times on gateway errors, which include timeouts after the server did the work. Make the mesh-wide default retry only connection failures and refused streams, allow broader retries per service for GET and HEAD, and use idempotency keys for writes.

Updated Houssam Hammoudi, CTOTested with Kuma 2.14.3 on Kubernetes 1.34.0 (kind)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Retries are a reliability feature. In a mesh they are also invisible: the sidecar next to the client repeats the request, and neither the client nor the server code knows it happened.

A retry is safe only if the first attempt had no effect, or if repeating it has the same effect (the request is idempotent). HTTP GET and HEAD are meant to be idempotent. POST usually is not: create an order, move money, send an email.

The dangerous case is a timeout or a 5xx after the server did the work. The server charged the card, then the response was slow, a gateway returned 504, or the connection dropped. The proxy sees a gateway error and retries. The server charges again.

What ships by default:

  • Kuma creates a default MeshRetry in every new mesh: HTTP requests are retried up to 5 times with a 16-second per-try timeout. With no retryOn set, Kuma uses gateway-error,connect-failure,refused-stream, and gateway-error covers 502, 503, 504 and "no response at all". Any method.
  • Istio retries HTTP requests twice by default, on connect-failure,refused-stream,unavailable,cancelled (set cluster-wide with meshConfig.defaultHttpRetryPolicy). That list does not include gateway-error or 5xx, so Istio's default does not retry a 504, but it applies to every method.

The security side is not only double charges. Retries repeat side effects an attacker can trigger (password reset mails, account creation), and they multiply load on a struggling backend: five retries in the mesh behind three in a client library behind two at the edge is up to (1+5) x (1+3) x (1+2) = 72 requests for one click.

What the docs say

This policy is similar to the 5xx policy but will attempt a retry if the upstream server responds with 502, 503, or 504 response code, or does not respond at all (disconnect/reset/read timeout).

Source: Envoy docs, Router filter, x-envoy-retry-on: gateway-error

Envoy will attempt a retry if the upstream server resets the stream with a REFUSED_STREAM error code. This reset type indicates that a request is safe to retry.

Source: Envoy docs, Router filter, x-envoy-retry-on: refused-stream

The default retry behavior for HTTP requests is to retry twice before returning the error.

Source: Istio docs, Traffic Management

For HTTP these are related to the response status code or method (5xx, 429, HttpMethodGet).

Source: Kuma docs, MeshRetry

Kuma's MeshRetry page lists the conditions but not the defaults a new mesh gets; those are in the source (HttpRetryOnDefault and the default mesh-retry-all policy). The Kuma page does not warn that its default retries non-idempotent methods after a 504 or a read timeout.

The secure configuration

1. Know what your mesh does today.

bash
kubectl get meshretries -A
kubectl -n kuma-system get meshretry -o yaml | grep -E 'name:|numRetries|retryOn' -A3

A mesh created with defaults has a policy named mesh-retry-all-<mesh> (mesh-retry-all-default for the default mesh) targeting the whole mesh.

2. A mesh-wide default that only retries what never reached a server.

yaml
apiVersion: kuma.io/v1alpha1
kind: MeshRetry
metadata:
  name: retry-safe-default
  namespace: kuma-system
  labels:
    kuma.io/mesh: default
spec:
  targetRef:
    kind: Mesh
  to:
    - targetRef:
        kind: Mesh
      default:
        http:
          numRetries: 2
          perTryTimeout: 5s
          backOff:
            baseInterval: 50ms
            maxInterval: 500ms
          retryOn:
            - ConnectFailure      # TCP connect failed: nothing was sent
            - RefusedStream       # HTTP/2 REFUSED_STREAM: server says it did not process it
        grpc:
          numRetries: 0           # off: no gRPC condition means "not processed"; enable per idempotent service

Then delete the built-in default policy, or create new meshes without it:

yaml
apiVersion: kuma.io/v1alpha1
kind: Mesh
metadata:
  name: default
spec:
  skipCreatingInitialPolicies:
    - MeshRetry

3. Broader retries only where the service owner says reads are safe. A producer policy in the service's namespace; the method conditions restrict every other condition in the list to GET and HEAD.

yaml
apiVersion: kuma.io/v1alpha1
kind: MeshRetry
metadata:
  name: catalog-read-retries
  namespace: catalog
  labels:
    kuma.io/mesh: default
spec:
  targetRef:
    kind: Mesh
  to:
    - targetRef:
        kind: MeshService
        name: catalog
        namespace: catalog
      default:
        http:
          numRetries: 3
          perTryTimeout: 2s
          retryOn:
            - GatewayError
            - ConnectFailure
            - RefusedStream
            - HttpMethodGet       # method conditions restrict every condition above to these methods
            - HttpMethodHead

4. Make writes safe to repeat in the application. Mesh settings reduce duplicates; they cannot remove them (a client, a user or a queue can still send twice). Accept an idempotency key on every state-changing endpoint, store it with the result, and return the stored result for a repeated key:

http
POST /v1/payments
Idempotency-Key: 7f3c2a4e-9d1b-4c55-8a6e-2b9f0c1d3e44

5. On Istio, set retries explicitly per route instead of relying on the default, and keep them to connection-level failures for writes:

yaml
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
  name: payments
  namespace: payments
spec:
  hosts:
    - payments.payments.svc.cluster.local
  http:
    - route:
        - destination:
            host: payments.payments.svc.cluster.local
      retries:
        attempts: 2
        perTryTimeout: 5s
        retryOn: connect-failure,refused-stream

Prove it

Run on a lab cluster (Kubernetes 1.34, Kuma 2.14.3). A test service in payments logs every request and holds POST /slow for 30 seconds; the mesh's request timeout is Kuma's default 15 seconds. Real output.

1. The retry policy the sidecar actually runs:

bash
kumactl inspect dataplane frontend-5d9c7b8f6-abcde.app --type=config-dump \
  | grep -E '"retry_on"|"num_retries"'

With a policy that retries on Kuma's default conditions:

text
"num_retries": 5,
"retry_on": "gateway-error,connect-failure,refused-stream",

With retry-safe-default from this page:

text
"num_retries": 2,
"retry_on": "connect-failure,refused-stream",

2. A slow POST is not repeated:

bash
kubectl -n app exec deploy/frontend -c app -- \
  curl -s -o /dev/null -w '%{http_code} after %{time_total}s\n' -X POST http://payments-test.payments.svc.cluster.local:8080/slow
kubectl -n payments logs -l app=payments-test --tail=-1 | grep -c 'POST /slow'

With gateway-error in retryOn and a 5 second per-try timeout, one click reached the server three times before the 15 second request timeout ended it:

text
504 after 15.009961s
3

With this page's policy:

text
504 after 5.017690s
1

Three identical POSTs to a payment endpoint is three charges. The client saw one 504 either way; only the server's log tells them apart.

3. Reads still recover. With catalog-read-retries applied and one of two catalog replicas deleted mid-test:

bash
for i in $(seq 1 20); do
  kubectl -n app exec deploy/frontend -c app -- \
    curl -s -o /dev/null -w '%{http_code}\n' http://catalog.catalog.svc.cluster.local:8080/items
done | sort | uniq -c
text
     20 200

Mistakes people make

Trusting the default because it says "gateway error"

A 504 is a gateway error, and it often means the server is still working on the request. Retrying it duplicates the work.

Retrying on 5xx for everything

5xx includes "no response at all", which is what a timeout looks like after the server committed. Restrict broad conditions to idempotent methods.

Per-try timeouts shorter than the work

A per-try timeout of 2 seconds on an endpoint that takes 3 turns every slow request into a retry. Set timeouts from real latency, per service.

Retries in three layers

Client library, mesh and edge gateway each retry, and the counts multiply. Pick one layer for retries, usually the mesh, and switch the others off or down.

Thinking the mesh setting is enough

Networks and users can still send a request twice. Idempotency keys in the application are the only complete fix for writes.

Checklist

  • The built-in default MeshRetry is removed or the mesh skips it.
  • The mesh-wide MeshRetry retries only ConnectFailure and RefusedStream.
  • gRPC retries are off by default and enabled per idempotent service.
  • Broader retries exist only as producer policies with HttpMethodGet and HttpMethodHead.
  • Per-try timeouts match the service's real latency.
  • Only one layer (client, mesh or edge) retries a given call.
  • State-changing endpoints accept and store an idempotency key.
  • A slow test POST reaches the server exactly once.

Retries are a promise that doing it again is harmless. Make the mesh keep that promise only where the service can.

H2-CSSE

Learn it on a live range

Production services in Go and Python, in Secure Software: a real host in your browser, and every objective checked on the machine.

Start free

The Dome

Want it run for you?

The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.

See the Dome