Multi-region edge PoPs on Kubernetes, the secure way
The second region went live in an afternoon. The sixth one could not get a certificate, the third still runs last month's WAF config, and all of them share one private key that someone copied around by hand.
The short answer
Treat every PoP as a copy of one Git definition with a small per-PoP overlay. Each PoP issues its own certificate with a PoP-unique name added, over DNS-01, so no key leaves the region. Run Envoy with several replicas, anti-affinity and a PodDisruptionBudget, steer clients with health-checked DNS, and encrypt the PoP-to-origin hop with a pinned CA.
On this page
What goes wrong
An edge point of presence (PoP) is a small cluster close to users that terminates TLS, runs the WAF and forwards to the origin. Running several is mostly about keeping them identical and independent. The failures come from the opposite.
Certificates are the first surprise. Every PoP serves the same hostnames, so every PoP asks the CA for the same set of names. Let's Encrypt allows five certificates per exact set of names per week, globally. The sixth PoP, or the third redeploy of one, is refused. HTTP-01 challenges fail too: with GeoDNS or anycast, the CA's validation request lands on whichever PoP is nearest to the CA, not the one that asked. The common workaround, one certificate and key copied to every region, turns one leaked node into a leaked key for the whole edge.
Drift is the second. A PoP configured by hand misses the next WAF rule, the next TLS change, the next ban list. Attackers only need the weakest PoP.
The third is hidden coupling: a PoP that calls a service in another region for every request (a shared rate limit store, a central auth check) fails when that region does, which defeats the point of having regions.
What the docs say
Up to 5 certificates can be issued per exact same set of identifiers every 7 days. This is a global limit, and all new order requests, regardless of which account submits them, count towards this limit.
Source: Let's Encrypt, Rate Limits
Local ServiceExternalTrafficPolicyLocal preserves the source IP of the traffic by routing only to endpoints on the same node as the traffic was received on (dropping the traffic if there are no local endpoints).
Source: Envoy Gateway API reference, ServiceExternalTrafficPolicy
EnvoyPDB allows to control the pod disruption budget of an Envoy Proxy.
Source: Envoy Gateway API reference, EnvoyProxyKubernetesProvider
None of these pages is about multi-region edges. The rate limit page does not mention that identical PoPs are a common way to hit the limit.
The secure configuration
One repository, a shared base and a thin overlay per PoP:
edge/
base/ # identical everywhere
gateway.yaml # Gateway, listeners
client-traffic.yaml # TLS 1.3, ML-KEM groups, client IP detection
waf.yaml # EnvoyExtensionPolicy (Coraza, pinned digest)
security.yaml # ban list, admin allowlists
backend-tls.yaml # BackendTLSPolicy to the origin
pops/
edge-1/kustomization.yaml # base + PoP name, PoP certificate
edge-2/kustomization.yamlPer PoP, the Envoy fleet (the Gateway points at it with
infrastructure.parametersRef):
apiVersion: gateway.envoyproxy.io/v1alpha1
kind: EnvoyProxy
metadata:
name: edge-pop
namespace: edge
spec:
provider:
type: Kubernetes
kubernetes:
envoyDeployment:
replicas: 3 # survive a node loss and a rollout
pod:
affinity:
podAntiAffinity: # never two Envoys on one node
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels:
gateway.envoyproxy.io/owning-gateway-name: edge
gateway.envoyproxy.io/owning-gateway-namespace: edge # not same-named Gateways elsewhere
topologyKey: kubernetes.io/hostname
envoyPDB:
minAvailable: 2 # node drains keep two serving
envoyService:
externalTrafficPolicy: Local # keep the client source addressPer PoP, its own certificate, with its own key, over DNS-01:
apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: www-example-com
namespace: edge
spec:
secretName: www-example-com-tls
dnsNames:
- www.example.com
- edge-1.example.com # PoP-unique name: a different set per PoP
issuerRef:
kind: ClusterIssuer
name: letsencrypt-dns01 # DNS-01: works whichever PoP the CA reaches
privateKey:
algorithm: ECDSA
size: 384
rotationPolicy: Always # the key is generated in this PoP and stays hereA PoP whose certificate has not been issued does not serve at all: Envoy Gateway creates no Envoy fleet for a Gateway whose certificate Secret is missing. A new PoP with a stuck DNS-01 order stays dark instead of serving without TLS, and health-checked DNS keeps it out of rotation.
Steering: publish the PoPs in DNS with health checks against a path that goes through Envoy to the origin, and a short TTL, so a PoP that cannot serve is removed. See GeoDNS with Knot and DNSSEC.
Independence: keep request-time dependencies inside the PoP. A global rate limit store per PoP counts per PoP, which is weaker than one shared store and far better than every PoP failing with one region.
Origin hop: every PoP reaches the origin over TLS with a pinned CA; see upstream TLS with a pinned CA.
Prove it
Two PoPs, edge-1 and edge-2, each a kind cluster with a control plane
and three workers, Cilium for LoadBalancer addresses, Envoy Gateway v1.9.1
and cert-manager v1.21.2. Both were deployed with kubectl apply -k from the
same repository:
edge/
base/ namespace, Gateway and HTTPRoute, the EnvoyProxy above,
ClientTrafficPolicy (TLS 1.3), Coraza WAF (pinned digest), origin stand-in
pops/edge-1/ ../../base + the Certificate above with edge-1.example.com
pops/edge-2/ ../../base + the Certificate above with edge-2.example.comThe letsencrypt-dns01 ClusterIssuer pointed at Pebble, the ACME test
server, which validated against a BIND server that cert-manager updated with
a TSIG key (see DNS-01 with RFC 2136 and TSIG).
1. Three Envoys per PoP, one per node, with a disruption budget and the client address kept:
$ kubectl get pods -n envoy-gateway-system -l gateway.envoyproxy.io/owning-gateway-name=edge -o wide
NAME READY STATUS NODE
envoy-edge-edge-de8ab67e-5b7ff867f5-brgn8 2/2 Running edge-1-worker2
envoy-edge-edge-de8ab67e-5b7ff867f5-sb5kb 2/2 Running edge-1-worker3
envoy-edge-edge-de8ab67e-5b7ff867f5-zdjfp 2/2 Running edge-1-worker
$ kubectl -n envoy-gateway-system get pdb
NAME MIN AVAILABLE MAX UNAVAILABLE ALLOWED DISRUPTIONS
envoy-edge-edge-de8ab67e 2 N/A 1
$ kubectl -n envoy-gateway-system get svc ... EXTERNAL-IP, externalTrafficPolicy
envoy-edge-edge-de8ab67e 172.18.252.10 LocalThe same in edge-2, on its own three workers. Envoy's access log for a
request from the test host (172.18.0.1):
{"method":"GET","response_code":200,"x-forwarded-for":"172.18.0.1","downstream_remote_address":"172.18.0.1:43886"}2. A certificate and a key per PoP:
kubectl get secret www-example-com-tls -n edge -o jsonpath='{.data.tls\.crt}' \
| base64 -d | openssl x509 -noout -ext subjectAltName
kubectl get secret www-example-com-tls -n edge -o jsonpath='{.data.tls\.crt}' \
| base64 -d | openssl x509 -noout -pubkey | sha256sumedge-1: DNS:www.example.com, DNS:edge-1.example.com public key sha256: 8fd04f212853a3d9...
edge-2: DNS:www.example.com, DNS:edge-2.example.com public key sha256: fbd4a19f392c48c3...3. Same policy result from every PoP, and a drifted PoP shows up. From outside, pinned to each PoP's address:
curl -s -o /dev/null -w "%{http_code}\n" --resolve www.example.com:443:192.0.2.10 \
"https://www.example.com/?id=1%27%20OR%20%271%27%3D%271"both PoPs as synced: edge-1 normal 200 sqli 403 edge-2 normal 200 sqli 403
WAF policy deleted on edge-2: edge-1 sqli 403 edge-2 sqli 200The 200 from one address is the drift. Run this check from outside after
every sync, against every PoP address, not against the load-balanced name.
The resources also pass offline checks with egctl v1.9.1
(secure-tests/multi-region-edge-pops-kubernetes/run.sh): schema validation
of the EnvoyProxy and Gateway, a caught typo (replicas: "three"), and a
translation of the full resource set.
Mistakes people make
One certificate, one key, everywhere
A copied key means every PoP is as safe as the least safe one, and rotating it is a fleet-wide event. Issue per PoP.
Identical name sets in every PoP
Six PoPs asking for the same names hit the weekly limit. Add a PoP-unique name to each certificate, or keep an eye on the count before you add PoPs.
HTTP-01 behind GeoDNS or anycast
The validation request reaches the nearest PoP, not the requesting one. Use DNS-01.
Hand-edited PoPs
Every manual change is a drift waiting to be found by an attacker. Change the base in Git and let every PoP sync.
Health checks that stop at the edge
A PoP that serves its own health page but cannot reach the origin still attracts traffic. Check a path that goes to the origin.
Two replicas on one node
Without anti-affinity, a node failure can take the whole PoP down. Require anti-affinity and a PodDisruptionBudget.
Checklist
- Define every PoP from one Git base with a thin per-PoP overlay.
- Issue a certificate per PoP with a PoP-unique name, over DNS-01.
- Keep private keys in their PoP; never copy them.
- Run three Envoy replicas with required anti-affinity and a PDB.
- Preserve the client address (
externalTrafficPolicy: Localor PROXY protocol). - Health-check PoPs through to the origin, with a short DNS TTL.
- Keep request-time dependencies inside each PoP.
- Encrypt the PoP-to-origin hop with a pinned CA.
- Test the same attack against every PoP address and compare.
Many PoPs, one definition. The day they stop being copies of each other is the day the weakest one becomes your edge.
H2-CPQE
Learn it on a live range
Edge points of presence, in Edge and Post-Quantum Networking: a real host in your browser, and every objective checked on the machine.
Start freeThe Dome
Want it run for you?
The Dome puts post-quantum TLS, a WAF that blocks, signed DNS and a zero-trust mesh in front of your application. Tell us what you run.
See the Dome