Retries and non-idempotent requests in a mesh, the secure way
A customer was charged three times, the payment service logs show three requests, and the frontend swears it sent one. Nobody wrote a retry loop. The mesh did, politely, because the first answer was a 504.
The short answer
A mesh proxy retries requests the application never asked to repeat. Kuma's default mesh retries HTTP up to five times on gateway errors, which include timeouts after the server did the work. Make the mesh-wide default retry only connection failures and refused streams, allow broader retries per service for GET and HEAD, and use idempotency keys for writes.
On this page
What goes wrong
Retries are a reliability feature. In a mesh they are also invisible: the sidecar next to the client repeats the request, and neither the client nor the server code knows it happened.
A retry is safe only if the first attempt had no effect, or if repeating it
has the same effect (the request is idempotent). HTTP GET and HEAD are
meant to be idempotent. POST usually is not: create an order, move money,
send an email.
The dangerous case is a timeout or a 5xx after the server did the work. The
server charged the card, then the response was slow, a gateway returned
504, or the connection dropped. The proxy sees a gateway error and retries.
The server charges again.
What ships by default:
- Kuma creates a default
MeshRetryin every new mesh: HTTP requests are retried up to 5 times with a 16-second per-try timeout. With noretryOnset, Kuma usesgateway-error,connect-failure,refused-stream, andgateway-errorcovers 502, 503, 504 and "no response at all". Any method. - Istio retries HTTP requests twice by default, on
connect-failure,refused-stream,unavailable,cancelled(set cluster-wide withmeshConfig.defaultHttpRetryPolicy). That list does not includegateway-erroror5xx, so Istio's default does not retry a 504, but it applies to every method.
The security side is not only double charges. Retries repeat side effects an attacker can trigger (password reset mails, account creation), and they multiply load on a struggling backend: five retries in the mesh behind three in a client library behind two at the edge is up to (1+5) x (1+3) x (1+2) = 72 requests for one click.
What the docs say
This policy is similar to the 5xx policy but will attempt a retry if the upstream server responds with 502, 503, or 504 response code, or does not respond at all (disconnect/reset/read timeout).
Source: Envoy docs, Router filter, x-envoy-retry-on: gateway-error
Envoy will attempt a retry if the upstream server resets the stream with a REFUSED_STREAM error code. This reset type indicates that a request is safe to retry.
Source: Envoy docs, Router filter, x-envoy-retry-on: refused-stream
The default retry behavior for HTTP requests is to retry twice before returning the error.
Source: Istio docs, Traffic Management
For HTTP these are related to the response status code or method (5xx, 429, HttpMethodGet).
Source: Kuma docs, MeshRetry
Kuma's MeshRetry page lists the conditions but not the defaults a new mesh
gets; those are in the source (HttpRetryOnDefault and the default
mesh-retry-all policy). The Kuma page does not warn that its default
retries non-idempotent methods after a 504 or a read timeout.
The secure configuration
1. Know what your mesh does today.
kubectl get meshretries -A
kubectl -n kuma-system get meshretry -o yaml | grep -E 'name:|numRetries|retryOn' -A3A mesh created with defaults has a policy named mesh-retry-all-<mesh>
(mesh-retry-all-default for the default mesh) targeting the whole mesh.
2. A mesh-wide default that only retries what never reached a server.
apiVersion: kuma.io/v1alpha1
kind: MeshRetry
metadata:
name: retry-safe-default
namespace: kuma-system
labels:
kuma.io/mesh: default
spec:
targetRef:
kind: Mesh
to:
- targetRef:
kind: Mesh
default:
http:
numRetries: 2
perTryTimeout: 5s
backOff:
baseInterval: 50ms
maxInterval: 500ms
retryOn:
- ConnectFailure # TCP connect failed: nothing was sent
- RefusedStream # HTTP/2 REFUSED_STREAM: server says it did not process it
grpc:
numRetries: 0 # off: no gRPC condition means "not processed"; enable per idempotent serviceThen delete the built-in default policy, or create new meshes without it:
apiVersion: kuma.io/v1alpha1
kind: Mesh
metadata:
name: default
spec:
skipCreatingInitialPolicies:
- MeshRetry3. Broader retries only where the service owner says reads are safe. A
producer policy in the service's namespace; the method conditions restrict
every other condition in the list to GET and HEAD.
apiVersion: kuma.io/v1alpha1
kind: MeshRetry
metadata:
name: catalog-read-retries
namespace: catalog
labels:
kuma.io/mesh: default
spec:
targetRef:
kind: Mesh
to:
- targetRef:
kind: MeshService
name: catalog
namespace: catalog
default:
http:
numRetries: 3
perTryTimeout: 2s
retryOn:
- GatewayError
- ConnectFailure
- RefusedStream
- HttpMethodGet # method conditions restrict every condition above to these methods
- HttpMethodHead4. Make writes safe to repeat in the application. Mesh settings reduce duplicates; they cannot remove them (a client, a user or a queue can still send twice). Accept an idempotency key on every state-changing endpoint, store it with the result, and return the stored result for a repeated key:
POST /v1/payments
Idempotency-Key: 7f3c2a4e-9d1b-4c55-8a6e-2b9f0c1d3e445. On Istio, set retries explicitly per route instead of relying on the default, and keep them to connection-level failures for writes:
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: payments
namespace: payments
spec:
hosts:
- payments.payments.svc.cluster.local
http:
- route:
- destination:
host: payments.payments.svc.cluster.local
retries:
attempts: 2
perTryTimeout: 5s
retryOn: connect-failure,refused-streamProve it
Run on a lab cluster (Kubernetes 1.34, Kuma 2.14.3). A test service in
payments logs every request and holds POST /slow for 30 seconds; the
mesh's request timeout is Kuma's default 15 seconds. Real output.
1. The retry policy the sidecar actually runs:
kumactl inspect dataplane frontend-5d9c7b8f6-abcde.app --type=config-dump \
| grep -E '"retry_on"|"num_retries"'With a policy that retries on Kuma's default conditions:
"num_retries": 5,
"retry_on": "gateway-error,connect-failure,refused-stream",With retry-safe-default from this page:
"num_retries": 2,
"retry_on": "connect-failure,refused-stream",2. A slow POST is not repeated:
kubectl -n app exec deploy/frontend -c app -- \
curl -s -o /dev/null -w '%{http_code} after %{time_total}s\n' -X POST http://payments-test.payments.svc.cluster.local:8080/slow
kubectl -n payments logs -l app=payments-test --tail=-1 | grep -c 'POST /slow'With gateway-error in retryOn and a 5 second per-try timeout, one click
reached the server three times before the 15 second request timeout ended it:
504 after 15.009961s
3With this page's policy:
504 after 5.017690s
1Three identical POSTs to a payment endpoint is three charges. The client saw
one 504 either way; only the server's log tells them apart.
3. Reads still recover. With catalog-read-retries applied and one of
two catalog replicas deleted mid-test:
for i in $(seq 1 20); do
kubectl -n app exec deploy/frontend -c app -- \
curl -s -o /dev/null -w '%{http_code}\n' http://catalog.catalog.svc.cluster.local:8080/items
done | sort | uniq -c 20 200Mistakes people make
Trusting the default because it says "gateway error"
A 504 is a gateway error, and it often means the server is still working on the request. Retrying it duplicates the work.
Retrying on 5xx for everything
5xx includes "no response at all", which is what a timeout looks like after
the server committed. Restrict broad conditions to idempotent methods.
Per-try timeouts shorter than the work
A per-try timeout of 2 seconds on an endpoint that takes 3 turns every slow request into a retry. Set timeouts from real latency, per service.
Retries in three layers
Client library, mesh and edge gateway each retry, and the counts multiply. Pick one layer for retries, usually the mesh, and switch the others off or down.
Thinking the mesh setting is enough
Networks and users can still send a request twice. Idempotency keys in the application are the only complete fix for writes.
Checklist
- The built-in default MeshRetry is removed or the mesh skips it.
- The mesh-wide MeshRetry retries only
ConnectFailureandRefusedStream. - gRPC retries are off by default and enabled per idempotent service.
- Broader retries exist only as producer policies with
HttpMethodGetandHttpMethodHead. - Per-try timeouts match the service's real latency.
- Only one layer (client, mesh or edge) retries a given call.
- State-changing endpoints accept and store an idempotency key.
- A slow test POST reaches the server exactly once.
Retries are a promise that doing it again is harmless. Make the mesh keep that promise only where the service can.
H2-CSSE
Learn it on a live range
Production services in Go and Python, in Secure Software: a real host in your browser, and every objective checked on the machine.
Start freeThe Dome
Want it run for you?
The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.
See the Dome