Nodes and clusters

etcd on cloud VMs, the secure way

Your cluster works until a neighbor on the same host decides to rebuild their search index. Then etcd misses a heartbeat, the leader changes three times in a minute, and kubectl answers every question with a timeout.

The short answer

Run etcd on the private network only, with client and peer certificate authentication and TLS 1.3. Give it a dedicated disk, check that WAL fsync p99 stays under 10 ms, and raise the heartbeat and election timeout together when the cloud is noisy. Keep metrics on localhost, and encrypt every snapshot before it leaves the node.

Updated Houssam Hammoudi, CTOTested with etcd 3.7.1, OpenSSL 3.5 (alpine:3.22), talosctl 1.14.1

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

etcd is the Kubernetes database. Every object lives there: Secrets, ServiceAccount tokens stored as Secrets, ConfigMaps, and every pod spec with its environment variables. Whoever talks to etcd directly skips Kubernetes authentication, RBAC, admission and the audit log.

etcd's own defaults are made for a laptop. Out of the box it serves plain HTTP, asks no client for a certificate, and has authentication off. Kubernetes installers turn TLS on, but it is easy to lose on the way: a second listener added "for debugging", peer traffic left on the public interface because the VM has one, or metrics exposed on the client port.

Cloud VMs add a second problem: timing. etcd writes every change to its log with fsync before it answers. Network block storage shared with other tenants can stall that write for hundreds of milliseconds. A stalled leader misses heartbeats, followers start an election, and the API server times out. People react by moving etcd to whatever disk is fastest, including the unencrypted local scratch disk.

The third problem is the backup. An etcd snapshot is a complete copy of the database. Unless Kubernetes encrypts Secrets at rest, they are in it in plain text, and everything else always is. Snapshots end up in object storage with the same protection as build logs.

What the docs say

Note that etcd doesn't enable RBAC based authentication or the authentication feature in the transport layer by default to reduce friction for users getting started with the database.

Source: etcd docs v3.7, Security model

No. etcd doesn't encrypt key/value data stored on disk drives.

Source: etcd docs v3.7, Transport security model (FAQ: Does etcd encrypt data stored on disk drives?)

monitor wal_fsync_duration_seconds (p99 duration should be less than 10ms) to confirm the disk is reasonably fast.

Source: etcd docs v3.7, FAQ

Election timeouts must be at least 10 times the round-trip time so it can account for variance in the network.

Source: etcd docs v3.7, Tuning

The tuning page talks about network round-trip time. On cloud VMs inside one region the round trip is about a millisecond; the stalls come from disks and CPU steal, which the same page mentions only in passing. None of the pages says that a snapshot is a plaintext copy of every Secret that Kubernetes did not encrypt itself.

The secure configuration

1. One member, as a script. The same flags work in a systemd unit or a static pod. The address is the VM's private (VPC) address.

sh
# etcd-member.sh
# One etcd member on a cloud VM: private address only, mTLS for clients and peers.
NAME="${NAME:-etcd-1}"
IP="${IP:-10.10.1.11}"   # private VPC address; never a public one
CLUSTER="${CLUSTER:-etcd-1=https://10.10.1.11:2380,etcd-2=https://10.10.1.12:2380,etcd-3=https://10.10.1.13:2380}"
PKI=/etc/etcd/pki
exec etcd --name "$NAME" --data-dir /var/lib/etcd \
  --listen-client-urls "https://$IP:2379" --advertise-client-urls "https://$IP:2379" \
  --listen-peer-urls "https://$IP:2380" --initial-advertise-peer-urls "https://$IP:2380" \
  --initial-cluster "$CLUSTER" --initial-cluster-state new \
  --client-cert-auth --trusted-ca-file "$PKI/ca.crt" \
  --cert-file "$PKI/server.crt" --key-file "$PKI/server.key" \
  --peer-client-cert-auth --peer-trusted-ca-file "$PKI/ca.crt" \
  --peer-cert-file "$PKI/server.crt" --peer-key-file "$PKI/server.key" \
  --tls-min-version TLS1.3 \
  --listen-metrics-urls http://127.0.0.1:2381 \
  --heartbeat-interval 250 --election-timeout 2500 \
  --quota-backend-bytes 8589934592 \
  --auto-compaction-mode periodic --auto-compaction-retention 1h

What each security line does:

  • --client-cert-auth and --peer-client-cert-auth with a dedicated etcd CA: only certificates signed by that CA connect. Use a CA that signs nothing else, so a Kubernetes client certificate cannot reach etcd.
  • --listen-*-urls on the private address: no listener on the public interface, and no plain http:// listener at all.
  • --tls-min-version TLS1.3: older protocol versions are refused.
  • --listen-metrics-urls http://127.0.0.1:2381: metrics and health on localhost, not on the client port where they would need a client certificate or, worse, get one.
  • --heartbeat-interval 250 --election-timeout 2500: the ratio stays 1:10. Raise them only after measuring; higher values mean slower failover.

2. The same on Talos. Talos generates the PKI and flags. Pin peer traffic to the private network, put etcd on its own encrypted volume, and set the timings:

yaml
# talos-etcd.yaml
# talosctl gen config ... --config-patch-control-plane @talos-etcd.yaml
cluster:
  etcd:
    advertisedSubnets:
      - 10.10.0.0/16            # etcd peers and clients use the VPC address, not the public one
    extraArgs:
      heartbeat-interval: "250"
      election-timeout: "2500"
      quota-backend-bytes: "8589934592"
---
apiVersion: v1alpha1
kind: VolumeConfig
name: ETCD                      # Talos 1.14+: etcd on a dedicated partition instead of /var
provisioning:
  diskSelector:
    match: disk.dev_path == "/dev/sdb"   # a dedicated volume attached to the VM
  minSize: 10GiB
  grow: false
encryption:
  provider: luks2
  keys:
    - slot: 0
      nodeID: {}
      lockToState: true         # only unlocks with this node's STATE (encrypt STATE too)

3. Firewall. Allow 2379 and 2380 only from the control plane nodes' private addresses, at the cloud firewall and on the host (see Talos Linux hardening and Cloud firewalls by tag).

4. Encrypted snapshots. Encrypt before the file leaves the node, to a key the cluster does not hold:

bash
etcdctl snapshot save /tmp/etcd.db   # on Talos: talosctl -n 10.10.1.11 etcd snapshot /tmp/etcd.db
etcdutl snapshot status /tmp/etcd.db -w table
age -r age1zcqwak5adwj8pdk9zykp5av6v0qyh2d5kkcqkcpf59d64r25saasgm5n0y \
  -o "etcd-$(date +%F).db.age" /tmp/etcd.db
shred -u /tmp/etcd.db

Keep Kubernetes Secrets encryption at rest on as well. Talos configures a secretbox provider for Secrets when it creates the cluster.

Prove it

Real output from secure-tests/etcd-cloud-vms/: one etcd 3.7.1 member in a container, started with the script above (IP=127.0.0.1), and a test CA.

1. No client certificate, no access:

bash
etcdctl --endpoints https://127.0.0.1:2379 --cacert ca.crt endpoint health
text
https://127.0.0.1:2379 is unhealthy: failed to commit proposal: context deadline exceeded
Error: unhealthy cluster

The etcd log says why:

text
rejected connection on client endpoint","remote-addr":"127.0.0.1:56718","server-name":"","error":"tls: client didn't provide a certificate"

2. A certificate from another CA, no access:

bash
etcdctl --endpoints https://127.0.0.1:2379 --cacert ca.crt \
  --cert other-ca-client.crt --key other-ca-client.key endpoint health
text
https://127.0.0.1:2379 is unhealthy: failed to commit proposal: context deadline exceeded
Error: unhealthy cluster
text
rejected connection on client endpoint","remote-addr":"127.0.0.1:56736","server-name":"","error":"tls: failed to verify certificate: x509: certificate signed by unknown authority"

3. Plain HTTP, no answer:

bash
curl -sS -m 3 http://127.0.0.1:2379/health
text
curl: (52) Empty reply from server

4. The API server's client certificate works:

bash
etcdctl --endpoints https://127.0.0.1:2379 --cacert ca.crt \
  --cert apiserver-etcd-client.crt --key apiserver-etcd-client.key endpoint health
text
https://127.0.0.1:2379 is healthy: successfully committed proposal: took = 7.857857ms

5. TLS 1.2 is refused:

bash
openssl s_client -connect 127.0.0.1:2379 -tls1_2 -CAfile ca.crt </dev/null
text
286B16B9627E0000:error:0A00042E:SSL routines:ssl3_read_bytes:tlsv1 alert protocol version:ssl/record/rec_layer_s3.c:918:SSL alert number 70

6. The cloud timings are in effect, and metrics answer on localhost only:

bash
curl -s http://127.0.0.1:2381/metrics | grep -E '^etcd_server_(has_leader|heartbeat_send_failures_total) '
text
etcd_server_has_leader 1
etcd_server_heartbeat_send_failures_total 0

The startup log confirms the values: "heartbeat-interval":"250ms","election-timeout":"2.5s".

7. A snapshot is a plaintext copy. Write a value, take a snapshot, and search the file:

bash
etcdctl put /registry/configmaps/app/db "password=hunter2"
etcdctl snapshot save /tmp/snap.db
grep -c hunter2 /tmp/snap.db
etcdutl snapshot status /tmp/snap.db -w table
text
Server version 3.7.0
1
┌──────────┬──────────┬────────────┬────────────┬─────────┐
│   HASH   │ REVISION │ TOTAL KEYS │ TOTAL SIZE │ VERSION │
├──────────┼──────────┼────────────┼────────────┼─────────┤
│ b0f53dc8 │        2 │          1 │      20 kB │   3.7.0 │
└──────────┴──────────┴────────────┴────────────┴─────────┘

The password is readable in the file. Server version 3.7.0 is the storage format version etcd 3.7.1 reports, not a different binary.

On a real cluster, check the disk with the metric the etcd FAQ names:

bash
curl -s http://127.0.0.1:2381/metrics | grep etcd_disk_wal_fsync_duration_seconds_bucket

What you should see: at least 99 percent of fsyncs in the buckets up to le="0.008" or le="0.016", and etcd_server_leader_changes_seen_total flat over a day.

Mistakes people make

Trusting the Kubernetes CA for etcd

If etcd trusts the same CA that signs kubelet and user certificates, every one of those certificates is an etcd client. Use a CA that signs only etcd members and the API server's etcd client certificate. kubeadm and Talos both do this; hand-built clusters often do not.

Listening on every interface

--listen-client-urls https://0.0.0.0:2379 on a VM with a public address puts the database on the internet, protected only by the client certificate check. Bind to the private address, and firewall it anyway.

Raising the timeouts instead of fixing the disk

A two-second election timeout hides a disk that stalls for a second. Measure wal_fsync_duration_seconds first. If p99 is above 10 ms, move etcd to a dedicated volume or a faster disk class, then tune.

Different timings on different members

The etcd tuning page says the heartbeat and election values must match on all members. Set them in one place (the Talos patch or one unit template).

Snapshots in the backup bucket, unencrypted

An etcd snapshot holds every object in the cluster. Encrypt it with a key the cluster cannot read, and store it where the cluster's own credentials cannot delete it.

Checklist

  • etcd listens only on private addresses, with https:// for clients and peers.
  • --client-cert-auth and --peer-client-cert-auth are on, with an etcd-only CA.
  • --tls-min-version TLS1.3 is set.
  • Metrics listen on 127.0.0.1:2381, not on the client port.
  • On Talos, cluster.etcd.advertisedSubnets names the private network.
  • Ports 2379 and 2380 are open only between control plane nodes.
  • WAL fsync p99 stays under 10 ms; etcd has a dedicated volume.
  • Heartbeat and election timeout are equal on all members, with a 1:10 ratio.
  • Kubernetes Secrets are encrypted at rest.
  • Snapshots are encrypted with age before leaving the node, and restores are tested.

etcd forgives slow networks and fast failovers, but not slow disks or open ports. Give it a quiet disk, a private address and its own CA, and it will stay boring, which is the best thing a database can be.

H2-CSPE

Learn it on a live range

Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on nodes and clusters

Talos Linux, Kubernetes API hardening, service account tokens, RBAC and the cloud underneath.

All nodes and clusters guides