Linux hosts

Kernel sysctl hardening, the secure way

You pasted forty sysctl lines from a hardening guide, ran sysctl --system, and felt safer. Two of those lines quietly cancel each other out, and one of them breaks every container on the host.

The short answer

Put hardening keys in one file under /etc/sysctl.d/, with a comment on each: kptr_restrict 2, dmesg_restrict 1, unprivileged BPF off, Yama ptrace_scope 2, protected_* file checks, and no redirects or source routing. Set ip_forward before the conf.* keys, because changing it resets them. Check the running values, not the file.

Updated Houssam Hammoudi, CTOTested with Linux 6.18 (WSL2 kernel), alpine:3.22 busybox sysctl

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

The kernel has hundreds of run-time settings under /proc/sys. Some defaults favor debugging and compatibility over security: kernel addresses visible to users, the kernel log readable by anyone, ptrace between a user's own processes, ICMP redirects accepted from the network.

Each of these helps a local attacker. Kernel addresses defeat address randomization. ptrace lets malware read secrets from another process of the same user, such as an SSH agent. Unprivileged BPF and userfaultfd have been parts of many kernel exploits.

Copy-pasted hardening lists cause the other problem. Order matters, some keys are one-way, and some break container hosts or routers. A file that looks right can leave the running kernel in a different state.

What the docs say

This variable is special, its change resets all configuration parameters to their default state (RFC1122 for hosts, RFC1812 for routers)

Source: Linux kernel docs, IP Sysctl, ip_forward

When kptr_restrict is set to 2, kernel pointers printed using %pK will be replaced with 0s regardless of privileges.

Source: Linux kernel docs, Documentation for /proc/sys/kernel/, kptr_restrict

Once set to 1, this can't be cleared from the running kernel anymore.

Source: Linux kernel docs, Documentation for /proc/sys/kernel/, unprivileged_bpf_disabled

The ip_forward sentence is easy to miss, and it is the one that bites. The kernel docs do not say what it means for a file of settings: ip_forward has to come first, or it undoes the redirect settings above it.

The secure configuration

conf
# /etc/sysctl.d/90-hardening.conf
# --- Kernel information leaks ---
# Hide kernel pointers in /proc and logs, even from root.
kernel.kptr_restrict = 2
# Only privileged users may read the kernel log (dmesg).
kernel.dmesg_restrict = 1
# Only privileged users may use perf; Debian and Ubuntu also accept 3 (no perf at all for users).
kernel.perf_event_paranoid = 3

# --- Attack surface for local exploits ---
# No BPF programs from unprivileged users; harden the BPF JIT.
kernel.unprivileged_bpf_disabled = 1
net.core.bpf_jit_harden = 2
# Only processes with CAP_SYS_PTRACE may attach to other processes.
kernel.yama.ptrace_scope = 2
# Do not allow loading a new kernel at run time.
kernel.kexec_load_disabled = 1
# No magic SysRq key combinations.
kernel.sysrq = 0
# Do not auto-load TTY line disciplines (a past source of kernel bugs).
dev.tty.ldisc_autoload = 0
# userfaultfd only for privileged users (used in many exploit chains).
vm.unprivileged_userfaultfd = 0
# Full address space layout randomization.
kernel.randomize_va_space = 2

# --- Filesystem traps in shared directories ---
fs.protected_symlinks = 1
fs.protected_hardlinks = 1
fs.protected_fifos = 2
fs.protected_regular = 2
# No core dumps from setuid programs (they may contain secrets).
fs.suid_dumpable = 0

# --- Network (per network namespace) ---
# FIRST: changing ip_forward resets accept_redirects and similar keys to defaults,
# so it must come before them. This host is not a router. (Container hosts and routers need 1.)
net.ipv4.ip_forward = 0
# Drop packets whose source address could not be routed back (anti-spoofing).
net.ipv4.conf.all.rp_filter = 1
net.ipv4.conf.default.rp_filter = 1
# Ignore ICMP redirects and source-routed packets.
net.ipv4.conf.all.accept_redirects = 0
net.ipv4.conf.default.accept_redirects = 0
net.ipv6.conf.all.accept_redirects = 0
net.ipv6.conf.default.accept_redirects = 0
net.ipv4.conf.all.accept_source_route = 0
net.ipv6.conf.all.accept_source_route = 0
# This host is not a router: send no redirects.
net.ipv4.conf.all.send_redirects = 0
# Log packets with impossible source addresses.
net.ipv4.conf.all.log_martians = 1
# SYN cookies under SYN flood.
net.ipv4.tcp_syncookies = 1
net.ipv4.icmp_echo_ignore_broadcasts = 1
bash
install -m 0644 90-hardening.conf /etc/sysctl.d/90-hardening.conf
sysctl --system            # applies every file in sysctl.d, in name order

Adjust for the host's role:

  • Container hosts and routers need net.ipv4.ip_forward = 1. Docker, Podman and Kubernetes nodes stop routing container traffic with 0.
  • Rootless containers need unprivileged user namespaces; do not add user.max_user_namespaces = 0 on those hosts.
  • Debuggers stop working for normal users with ptrace_scope = 2. Use 1 on developer machines.
  • kernel.unprivileged_bpf_disabled = 1 cannot be undone until reboot. 2 also blocks unprivileged BPF but lets an admin change it later.

A small script compares the file with the running kernel:

sh
#!/bin/sh
# Compare a sysctl.d file with the running kernel. Prints OK, DIFF or MISSING per key.
file=${1:-/etc/sysctl.d/90-hardening.conf}
grep -E '^[a-z]' "$file" | while IFS='=' read -r key want; do
  key=$(echo "$key" | tr -d ' '); want=$(echo "$want" | tr -d ' ')
  have=$(sysctl -n "$key" 2>/dev/null) || { printf '%-8s %-40s want %s\n' MISSING "$key" "$want"; continue; }
  if [ "$have" = "$want" ]; then printf '%-8s %-40s %s\n' OK "$key" "$have"; else printf '%-8s %-40s want %s, have %s\n' DIFF "$key" "$want" "$have"; fi
done

Prove it

The trap, shown in a container's own network namespace. Setting accept_redirects = 0 alone works. Setting it and then ip_forward (Docker applied ip_forward after the conf.* key here) silently brings redirects back:

bash
docker run --rm --sysctl net.ipv4.conf.all.accept_redirects=0 alpine:3.22 \
  sh -c "printf 'accept_redirects=0 alone:              '; sysctl -n net.ipv4.conf.all.accept_redirects"
docker run --rm --sysctl net.ipv4.conf.all.accept_redirects=0 --sysctl net.ipv4.ip_forward=0 alpine:3.22 \
  sh -c "printf 'accept_redirects=0, then ip_forward=0: '; sysctl -n net.ipv4.conf.all.accept_redirects"
text
accept_redirects=0 alone:              0
accept_redirects=0, then ip_forward=0: 1

Inside a container, /proc/sys is read-only, so sysctl -p cannot change kernel-wide keys (and must not; they would change the host):

bash
sysctl -p /t/90-hardening.conf 2>&1 | head -3
text
sysctl: error setting key 'kernel.kptr_restrict': Read-only file system
sysctl: error setting key 'kernel.dmesg_restrict': Read-only file system
sysctl: error setting key 'kernel.perf_event_paranoid': Read-only file system

The check script in a container that received the network keys with --sysctl (all except ip_forward). The kernel-wide values are the test machine's own kernel, which was not hardened:

bash
sh check.sh /t/90-hardening.conf
text
DIFF     kernel.kptr_restrict                     want 2, have 1
DIFF     kernel.dmesg_restrict                    want 1, have 0
DIFF     kernel.perf_event_paranoid               want 3, have 2
DIFF     kernel.unprivileged_bpf_disabled         want 1, have 2
MISSING  net.core.bpf_jit_harden                  want 2
DIFF     kernel.yama.ptrace_scope                 want 2, have 1
DIFF     kernel.kexec_load_disabled               want 1, have 0
DIFF     kernel.sysrq                             want 0, have 176
OK       dev.tty.ldisc_autoload                   0
OK       vm.unprivileged_userfaultfd              0
OK       kernel.randomize_va_space                2
OK       fs.protected_symlinks                    1
OK       fs.protected_hardlinks                   1
DIFF     fs.protected_fifos                       want 2, have 1
OK       fs.protected_regular                     2
OK       fs.suid_dumpable                         0
DIFF     net.ipv4.ip_forward                      want 0, have 1
OK       net.ipv4.conf.all.rp_filter              1
OK       net.ipv4.conf.default.rp_filter          1
OK       net.ipv4.conf.all.accept_redirects       0
OK       net.ipv4.conf.default.accept_redirects   0
OK       net.ipv6.conf.all.accept_redirects       0
OK       net.ipv6.conf.default.accept_redirects   0
OK       net.ipv4.conf.all.accept_source_route    0
OK       net.ipv6.conf.all.accept_source_route    0
OK       net.ipv4.conf.all.send_redirects         0
OK       net.ipv4.conf.all.log_martians           1
OK       net.ipv4.tcp_syncookies                  1
OK       net.ipv4.icmp_echo_ignore_broadcasts     1

How to read it: every network key applied (OK) except ip_forward, which the container inherits as 1 from its Docker host, exactly the container-host case above. net.core.bpf_jit_harden is MISSING because it exists only in the host's first network namespace. unprivileged_bpf_disabled is 2, which already blocks unprivileged BPF. The other DIFF lines are what this file would change on that kernel. The test is secure-tests/kernel-sysctl-hardening/run.sh.

Mistakes people make

Trusting the file instead of the kernel

A later file in /etc/sysctl.d/, a package, or a container runtime can set the same key. Compare with sysctl -n <key> or the check script after every boot.

ip_forward at the bottom

It resets accept_redirects, send_redirects and related keys to their defaults. Put it first in the network section.

Disabling forwarding on a container host

Containers lose network access. Keep ip_forward = 1 there, and filter forwarded traffic with the firewall instead.

Keys that do not exist

Old guides include keys removed from current kernels (for example net.ipv4.tcp_tw_recycle). sysctl --system prints an error for them and continues. Read the output.

One-way switches in testing

unprivileged_bpf_disabled = 1 and kexec_load_disabled = 1 cannot be turned off without a reboot. Try them on a test host first.

Checklist

  • One file in /etc/sysctl.d/ holds the hardening keys, with a comment per key.
  • net.ipv4.ip_forward comes before the conf.* keys.
  • The file matches the host's role (container host, router, developer machine).
  • sysctl --system runs without errors.
  • The check script shows OK for every key after a reboot.
  • kernel.kptr_restrict = 2 and kernel.dmesg_restrict = 1 are in force.
  • Unprivileged BPF is off (1 or 2).
  • kernel.yama.ptrace_scope is 1 or higher.
  • Redirects and source routing are off.

Sysctl hardening is forty small locks. The job is not to add more of them; it is to make sure each one is actually closed when the host is running.

FND

Learn it on a live range

Linux 3: securing Linux, in Foundation: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on linux hosts

SSH, sudo, firewalls, mount options, updates and the host basics every engineer should get right.

All linux hosts guides