systemd service sandboxing, the secure way
Your web app needs to read one config file and write to one directory. By default systemd lets it read the whole disk, write most of it, load kernel modules if it is root, and open raw sockets. That is a lot of trust for a program that parses strangers' input all day.
The short answer
Add sandboxing directives to each service unit: DynamicUser or a dedicated User, ProtectSystem=strict with StateDirectory for writable paths, ProtectHome, PrivateTmp, PrivateDevices, NoNewPrivileges, an empty CapabilityBoundingSet, RestrictAddressFamilies and SystemCallFilter=@system-service. Score the unit with systemd-analyze security and fail CI above a threshold.
On this page
What goes wrong
A service unit with only ExecStart= and User= runs with every right that
user has. It can read /home, /etc and anything else that is world-readable.
It can write to /tmp and /var/tmp shared with every other service. If it
runs as root, it also has every capability: loading kernel modules, changing
the clock, reading any file.
When an attacker finds a bug in that service, all of that becomes theirs. Sandboxing does not fix the bug. It shrinks what the bug can reach.
systemd already has the tools: private mounts, capability limits, syscall filters, network family limits. They are off by default, because turning them on can break a service that needs them.
What the docs say
If set to "strict" the entire file system hierarchy is mounted read-only, except for the API file system subtrees /dev/, /proc/ and /sys/ (protect these directories using PrivateDevices=, ProtectKernelTunables=, ProtectControlGroups=).
Source: systemd.exec(5), ProtectSystem=
If true, ensures that the service process and all its children can never gain new privileges through execve() (e.g. via setuid or setgid bits, or filesystem capabilities).
Source: systemd.exec(5), NoNewPrivileges=
Note that this only analyzes the per-service security features systemd itself implements.
Source: systemd-analyze(1), security
The score is a guide, not a verdict. A low exposure number means systemd's walls are up. It says nothing about whether the program inside has bugs.
The secure configuration
# /etc/systemd/system/app.service
[Unit]
Description=Example web app (sandboxed)
[Service]
ExecStart=/usr/local/bin/app --listen 127.0.0.1:8080
Restart=on-failure
# Identity: a throwaway user created at start, no fixed account on disk.
DynamicUser=yes
# Writable state lives only here (/var/lib/app, owned by the dynamic user).
StateDirectory=app
LogsDirectory=app
UMask=0027
# Filesystem: everything read-only except the directories above.
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectKernelLogs=yes
ProtectControlGroups=yes
ProtectClock=yes
ProtectHostname=yes
ProtectProc=invisible
ProcSubset=pid
# Privileges: no setuid escalation, no capabilities at all.
NoNewPrivileges=yes
CapabilityBoundingSet=
AmbientCapabilities=
RestrictSUIDSGID=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
RestrictRealtime=yes
RestrictNamespaces=yes
RemoveIPC=yes
PrivateUsers=yes
# Network: IPv4/IPv6 and local sockets only; no raw packets, no netlink.
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
IPAddressDeny=any
IPAddressAllow=localhost
# System calls: the common service set, minus privileged calls.
SystemCallArchitectures=native
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources
SystemCallErrorNumber=EPERM
DevicePolicy=closed
[Install]
WantedBy=multi-user.targetSome lines depend on the program. MemoryDenyWriteExecute=yes breaks
programs with a JIT (Java, Node.js, some Python extensions).
IPAddressAllow=localhost fits a service behind a local reverse proxy; list
the real peer addresses otherwise. For a packaged service, put the
directives in a drop-in (systemctl edit nginx) instead of editing the
vendor file.
Add one directive at a time on an existing service, restart, and watch
journalctl -u app for Permission denied, Read-only file system or
Operation not permitted.
Prove it
systemd-analyze verify reported no problem with the unit (the test filters
out the complaint that /usr/local/bin/app does not exist in the
container). The exposure scores, before and after:
systemd-analyze security --offline=true --no-pager app-plain.service | tail -1
systemd-analyze security --offline=true --no-pager app.service | tail -1→ Overall exposure level for app-plain.service: 9.2 UNSAFE :-{
→ Overall exposure level for app.service: 1.1 OK :-)What is left, and why each item stays:
systemd-analyze security --offline=true --no-pager app.service | grep "^✗"✗ RootDirectory=/RootImage= Service runs within the host's root directory 0.1
✗ PrivateMounts= Service may install system mounts 0.2
✗ RestrictAddressFamilies=~AF_UNIX Service may allocate local sockets 0.1
✗ RestrictAddressFamilies=~AF_(INET|INET6) Service may allocate Internet sockets 0.3
✗ PrivateNetwork= Service has access to the host's network 0.5
✗ DeviceAllow= Service has a device ACL with some special devices: char-rtc:r 0.1
✗ IPAddressDeny= Service defines IP address allow list with only localhost entries 0.1
✗ UMask= Files created by service are group-readable by default 0.1A web app needs Internet sockets and the host network, so those findings are
expected. --threshold turns the score into a CI gate: the command exits
non-zero when the exposure is above the number (here 30, meaning 3.0):
systemd-analyze security --offline=true --threshold=30 app-plain.service >/dev/null; echo "exit $?"
systemd-analyze security --offline=true --threshold=30 app.service >/dev/null; echo "exit $?"app-plain.service, threshold 30: exit 1
app.service, threshold 30: exit 0On a real host, check that the walls hold once the service runs:
systemctl show app -p User -p ProtectSystem -p NoNewPrivileges
cat /proc/$(systemctl show -p MainPID --value app)/status | grep -E "^(NoNewPrivs|CapEff|Seccomp)"
systemd-run --wait --pipe -p ProtectSystem=strict -p DynamicUser=yes touch /etc/probeWhat you should see: NoNewPrivs: 1, CapEff: 0000000000000000 and
Seccomp: 2 for the main process, and touch: cannot touch '/etc/probe': Read-only file system from the one-off systemd-run. The test script is
secure-tests/systemd-service-sandboxing/run.sh.
Mistakes people make
Hardening by copy-paste
A long block copied from a blog can break a service in ways that show up only under load (a JIT, a DNS lookup over netlink, a helper binary). Add directives in groups and test each group.
ProtectSystem=strict without a writable path
The service then fails to write its state or logs. Use StateDirectory=,
LogsDirectory=, CacheDirectory= or ReadWritePaths= for the paths it
really needs.
Leaving root services unsandboxed
Services that must start as root (to bind port 443, for example) benefit the
most. Use AmbientCapabilities=CAP_NET_BIND_SERVICE with a non-root user
instead of running as root.
Treating the score as the goal
Getting from 9.2 to 1.1 is useful. Getting from 1.1 to 0.9 by breaking the service's network is not. Stop when the remaining findings are things the service really needs.
Checklist
- Every custom service runs as a dedicated user or with
DynamicUser=yes. ProtectSystem=strictandProtectHome=yesare set, with explicit writable paths.PrivateTmp=yesandPrivateDevices=yesare set.NoNewPrivileges=yesis set andCapabilityBoundingSet=lists only what is needed.RestrictAddressFamilies=andSystemCallFilter=@system-serviceare set.systemd-analyze security <unit>scores each custom service, and the scores are recorded.- CI runs
systemd-analyze security --offline=true --threshold=...on unit files. - Vendor units are hardened with drop-ins, not by editing the vendor file.
A sandbox is a promise about the worst day. On that day, you will be glad the service could only reach the one directory it needed.
FND
Learn it on a live range
Linux 2: the kernel, storage, networks and the services a host runs, in Foundation: a real host in your browser, and every objective checked on the machine.
Start freeThe Secure Way
More on linux hosts
SSH, sudo, firewalls, mount options, updates and the host basics every engineer should get right.
All linux hosts guides