etcd on cloud VMs, the secure way
Your cluster works until a neighbor on the same host decides to rebuild their search index. Then etcd misses a heartbeat, the leader changes three times in a minute, and kubectl answers every question with a timeout.
The short answer
Run etcd on the private network only, with client and peer certificate authentication and TLS 1.3. Give it a dedicated disk, check that WAL fsync p99 stays under 10 ms, and raise the heartbeat and election timeout together when the cloud is noisy. Keep metrics on localhost, and encrypt every snapshot before it leaves the node.
On this page
What goes wrong
etcd is the Kubernetes database. Every object lives there: Secrets, ServiceAccount tokens stored as Secrets, ConfigMaps, and every pod spec with its environment variables. Whoever talks to etcd directly skips Kubernetes authentication, RBAC, admission and the audit log.
etcd's own defaults are made for a laptop. Out of the box it serves plain HTTP, asks no client for a certificate, and has authentication off. Kubernetes installers turn TLS on, but it is easy to lose on the way: a second listener added "for debugging", peer traffic left on the public interface because the VM has one, or metrics exposed on the client port.
Cloud VMs add a second problem: timing. etcd writes every change to its log
with fsync before it answers. Network block storage shared with other
tenants can stall that write for hundreds of milliseconds. A stalled leader
misses heartbeats, followers start an election, and the API server times out.
People react by moving etcd to whatever disk is fastest, including the
unencrypted local scratch disk.
The third problem is the backup. An etcd snapshot is a complete copy of the database. Unless Kubernetes encrypts Secrets at rest, they are in it in plain text, and everything else always is. Snapshots end up in object storage with the same protection as build logs.
What the docs say
Note that etcd doesn't enable RBAC based authentication or the authentication feature in the transport layer by default to reduce friction for users getting started with the database.
Source: etcd docs v3.7, Security model
No. etcd doesn't encrypt key/value data stored on disk drives.
Source: etcd docs v3.7, Transport security model (FAQ: Does etcd encrypt data stored on disk drives?)
monitor wal_fsync_duration_seconds (p99 duration should be less than 10ms) to confirm the disk is reasonably fast.
Source: etcd docs v3.7, FAQ
Election timeouts must be at least 10 times the round-trip time so it can account for variance in the network.
Source: etcd docs v3.7, Tuning
The tuning page talks about network round-trip time. On cloud VMs inside one region the round trip is about a millisecond; the stalls come from disks and CPU steal, which the same page mentions only in passing. None of the pages says that a snapshot is a plaintext copy of every Secret that Kubernetes did not encrypt itself.
The secure configuration
1. One member, as a script. The same flags work in a systemd unit or a static pod. The address is the VM's private (VPC) address.
# etcd-member.sh
# One etcd member on a cloud VM: private address only, mTLS for clients and peers.
NAME="${NAME:-etcd-1}"
IP="${IP:-10.10.1.11}" # private VPC address; never a public one
CLUSTER="${CLUSTER:-etcd-1=https://10.10.1.11:2380,etcd-2=https://10.10.1.12:2380,etcd-3=https://10.10.1.13:2380}"
PKI=/etc/etcd/pki
exec etcd --name "$NAME" --data-dir /var/lib/etcd \
--listen-client-urls "https://$IP:2379" --advertise-client-urls "https://$IP:2379" \
--listen-peer-urls "https://$IP:2380" --initial-advertise-peer-urls "https://$IP:2380" \
--initial-cluster "$CLUSTER" --initial-cluster-state new \
--client-cert-auth --trusted-ca-file "$PKI/ca.crt" \
--cert-file "$PKI/server.crt" --key-file "$PKI/server.key" \
--peer-client-cert-auth --peer-trusted-ca-file "$PKI/ca.crt" \
--peer-cert-file "$PKI/server.crt" --peer-key-file "$PKI/server.key" \
--tls-min-version TLS1.3 \
--listen-metrics-urls http://127.0.0.1:2381 \
--heartbeat-interval 250 --election-timeout 2500 \
--quota-backend-bytes 8589934592 \
--auto-compaction-mode periodic --auto-compaction-retention 1hWhat each security line does:
--client-cert-authand--peer-client-cert-authwith a dedicated etcd CA: only certificates signed by that CA connect. Use a CA that signs nothing else, so a Kubernetes client certificate cannot reach etcd.--listen-*-urlson the private address: no listener on the public interface, and no plainhttp://listener at all.--tls-min-version TLS1.3: older protocol versions are refused.--listen-metrics-urls http://127.0.0.1:2381: metrics and health on localhost, not on the client port where they would need a client certificate or, worse, get one.--heartbeat-interval 250 --election-timeout 2500: the ratio stays 1:10. Raise them only after measuring; higher values mean slower failover.
2. The same on Talos. Talos generates the PKI and flags. Pin peer traffic to the private network, put etcd on its own encrypted volume, and set the timings:
# talos-etcd.yaml
# talosctl gen config ... --config-patch-control-plane @talos-etcd.yaml
cluster:
etcd:
advertisedSubnets:
- 10.10.0.0/16 # etcd peers and clients use the VPC address, not the public one
extraArgs:
heartbeat-interval: "250"
election-timeout: "2500"
quota-backend-bytes: "8589934592"
---
apiVersion: v1alpha1
kind: VolumeConfig
name: ETCD # Talos 1.14+: etcd on a dedicated partition instead of /var
provisioning:
diskSelector:
match: disk.dev_path == "/dev/sdb" # a dedicated volume attached to the VM
minSize: 10GiB
grow: false
encryption:
provider: luks2
keys:
- slot: 0
nodeID: {}
lockToState: true # only unlocks with this node's STATE (encrypt STATE too)3. Firewall. Allow 2379 and 2380 only from the control plane nodes' private addresses, at the cloud firewall and on the host (see Talos Linux hardening and Cloud firewalls by tag).
4. Encrypted snapshots. Encrypt before the file leaves the node, to a key the cluster does not hold:
etcdctl snapshot save /tmp/etcd.db # on Talos: talosctl -n 10.10.1.11 etcd snapshot /tmp/etcd.db
etcdutl snapshot status /tmp/etcd.db -w table
age -r age1zcqwak5adwj8pdk9zykp5av6v0qyh2d5kkcqkcpf59d64r25saasgm5n0y \
-o "etcd-$(date +%F).db.age" /tmp/etcd.db
shred -u /tmp/etcd.dbKeep Kubernetes Secrets encryption at rest on as well. Talos configures a
secretbox provider for Secrets when it creates the cluster.
Prove it
Real output from secure-tests/etcd-cloud-vms/: one etcd 3.7.1 member in a
container, started with the script above (IP=127.0.0.1), and a test CA.
1. No client certificate, no access:
etcdctl --endpoints https://127.0.0.1:2379 --cacert ca.crt endpoint healthhttps://127.0.0.1:2379 is unhealthy: failed to commit proposal: context deadline exceeded
Error: unhealthy clusterThe etcd log says why:
rejected connection on client endpoint","remote-addr":"127.0.0.1:56718","server-name":"","error":"tls: client didn't provide a certificate"2. A certificate from another CA, no access:
etcdctl --endpoints https://127.0.0.1:2379 --cacert ca.crt \
--cert other-ca-client.crt --key other-ca-client.key endpoint healthhttps://127.0.0.1:2379 is unhealthy: failed to commit proposal: context deadline exceeded
Error: unhealthy clusterrejected connection on client endpoint","remote-addr":"127.0.0.1:56736","server-name":"","error":"tls: failed to verify certificate: x509: certificate signed by unknown authority"3. Plain HTTP, no answer:
curl -sS -m 3 http://127.0.0.1:2379/healthcurl: (52) Empty reply from server4. The API server's client certificate works:
etcdctl --endpoints https://127.0.0.1:2379 --cacert ca.crt \
--cert apiserver-etcd-client.crt --key apiserver-etcd-client.key endpoint healthhttps://127.0.0.1:2379 is healthy: successfully committed proposal: took = 7.857857ms5. TLS 1.2 is refused:
openssl s_client -connect 127.0.0.1:2379 -tls1_2 -CAfile ca.crt </dev/null286B16B9627E0000:error:0A00042E:SSL routines:ssl3_read_bytes:tlsv1 alert protocol version:ssl/record/rec_layer_s3.c:918:SSL alert number 706. The cloud timings are in effect, and metrics answer on localhost only:
curl -s http://127.0.0.1:2381/metrics | grep -E '^etcd_server_(has_leader|heartbeat_send_failures_total) 'etcd_server_has_leader 1
etcd_server_heartbeat_send_failures_total 0The startup log confirms the values: "heartbeat-interval":"250ms","election-timeout":"2.5s".
7. A snapshot is a plaintext copy. Write a value, take a snapshot, and search the file:
etcdctl put /registry/configmaps/app/db "password=hunter2"
etcdctl snapshot save /tmp/snap.db
grep -c hunter2 /tmp/snap.db
etcdutl snapshot status /tmp/snap.db -w tableServer version 3.7.0
1
┌──────────┬──────────┬────────────┬────────────┬─────────┐
│ HASH │ REVISION │ TOTAL KEYS │ TOTAL SIZE │ VERSION │
├──────────┼──────────┼────────────┼────────────┼─────────┤
│ b0f53dc8 │ 2 │ 1 │ 20 kB │ 3.7.0 │
└──────────┴──────────┴────────────┴────────────┴─────────┘The password is readable in the file. Server version 3.7.0 is the storage
format version etcd 3.7.1 reports, not a different binary.
On a real cluster, check the disk with the metric the etcd FAQ names:
curl -s http://127.0.0.1:2381/metrics | grep etcd_disk_wal_fsync_duration_seconds_bucketWhat you should see: at least 99 percent of fsyncs in the buckets up to
le="0.008" or le="0.016", and etcd_server_leader_changes_seen_total
flat over a day.
Mistakes people make
Trusting the Kubernetes CA for etcd
If etcd trusts the same CA that signs kubelet and user certificates, every one of those certificates is an etcd client. Use a CA that signs only etcd members and the API server's etcd client certificate. kubeadm and Talos both do this; hand-built clusters often do not.
Listening on every interface
--listen-client-urls https://0.0.0.0:2379 on a VM with a public address
puts the database on the internet, protected only by the client certificate
check. Bind to the private address, and firewall it anyway.
Raising the timeouts instead of fixing the disk
A two-second election timeout hides a disk that stalls for a second. Measure
wal_fsync_duration_seconds first. If p99 is above 10 ms, move etcd to a
dedicated volume or a faster disk class, then tune.
Different timings on different members
The etcd tuning page says the heartbeat and election values must match on all members. Set them in one place (the Talos patch or one unit template).
Snapshots in the backup bucket, unencrypted
An etcd snapshot holds every object in the cluster. Encrypt it with a key the cluster cannot read, and store it where the cluster's own credentials cannot delete it.
Checklist
- etcd listens only on private addresses, with
https://for clients and peers. --client-cert-authand--peer-client-cert-authare on, with an etcd-only CA.--tls-min-version TLS1.3is set.- Metrics listen on
127.0.0.1:2381, not on the client port. - On Talos,
cluster.etcd.advertisedSubnetsnames the private network. - Ports 2379 and 2380 are open only between control plane nodes.
- WAL fsync p99 stays under 10 ms; etcd has a dedicated volume.
- Heartbeat and election timeout are equal on all members, with a 1:10 ratio.
- Kubernetes Secrets are encrypted at rest.
- Snapshots are encrypted with age before leaving the node, and restores are tested.
etcd forgives slow networks and fast failovers, but not slow disks or open ports. Give it a quiet disk, a private address and its own CA, and it will stay boring, which is the best thing a database can be.
H2-CSPE
Learn it on a live range
Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on nodes and clusters
Talos Linux, Kubernetes API hardening, service account tokens, RBAC and the cloud underneath.
All nodes and clusters guides