Linux hosts

systemd service sandboxing, the secure way

Your web app needs to read one config file and write to one directory. By default systemd lets it read the whole disk, write most of it, load kernel modules if it is root, and open raw sockets. That is a lot of trust for a program that parses strangers' input all day.

The short answer

Add sandboxing directives to each service unit: DynamicUser or a dedicated User, ProtectSystem=strict with StateDirectory for writable paths, ProtectHome, PrivateTmp, PrivateDevices, NoNewPrivileges, an empty CapabilityBoundingSet, RestrictAddressFamilies and SystemCallFilter=@system-service. Score the unit with systemd-analyze security and fail CI above a threshold.

Updated Houssam Hammoudi, CTOTested with systemd 252.39 (Debian 12)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A service unit with only ExecStart= and User= runs with every right that user has. It can read /home, /etc and anything else that is world-readable. It can write to /tmp and /var/tmp shared with every other service. If it runs as root, it also has every capability: loading kernel modules, changing the clock, reading any file.

When an attacker finds a bug in that service, all of that becomes theirs. Sandboxing does not fix the bug. It shrinks what the bug can reach.

systemd already has the tools: private mounts, capability limits, syscall filters, network family limits. They are off by default, because turning them on can break a service that needs them.

What the docs say

If set to "strict" the entire file system hierarchy is mounted read-only, except for the API file system subtrees /dev/, /proc/ and /sys/ (protect these directories using PrivateDevices=, ProtectKernelTunables=, ProtectControlGroups=).

Source: systemd.exec(5), ProtectSystem=

If true, ensures that the service process and all its children can never gain new privileges through execve() (e.g. via setuid or setgid bits, or filesystem capabilities).

Source: systemd.exec(5), NoNewPrivileges=

Note that this only analyzes the per-service security features systemd itself implements.

Source: systemd-analyze(1), security

The score is a guide, not a verdict. A low exposure number means systemd's walls are up. It says nothing about whether the program inside has bugs.

The secure configuration

ini
# /etc/systemd/system/app.service
[Unit]
Description=Example web app (sandboxed)

[Service]
ExecStart=/usr/local/bin/app --listen 127.0.0.1:8080
Restart=on-failure

# Identity: a throwaway user created at start, no fixed account on disk.
DynamicUser=yes
# Writable state lives only here (/var/lib/app, owned by the dynamic user).
StateDirectory=app
LogsDirectory=app
UMask=0027

# Filesystem: everything read-only except the directories above.
ProtectSystem=strict
ProtectHome=yes
PrivateTmp=yes
PrivateDevices=yes
ProtectKernelTunables=yes
ProtectKernelModules=yes
ProtectKernelLogs=yes
ProtectControlGroups=yes
ProtectClock=yes
ProtectHostname=yes
ProtectProc=invisible
ProcSubset=pid

# Privileges: no setuid escalation, no capabilities at all.
NoNewPrivileges=yes
CapabilityBoundingSet=
AmbientCapabilities=
RestrictSUIDSGID=yes
LockPersonality=yes
MemoryDenyWriteExecute=yes
RestrictRealtime=yes
RestrictNamespaces=yes
RemoveIPC=yes
PrivateUsers=yes

# Network: IPv4/IPv6 and local sockets only; no raw packets, no netlink.
RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
IPAddressDeny=any
IPAddressAllow=localhost

# System calls: the common service set, minus privileged calls.
SystemCallArchitectures=native
SystemCallFilter=@system-service
SystemCallFilter=~@privileged @resources
SystemCallErrorNumber=EPERM
DevicePolicy=closed

[Install]
WantedBy=multi-user.target

Some lines depend on the program. MemoryDenyWriteExecute=yes breaks programs with a JIT (Java, Node.js, some Python extensions). IPAddressAllow=localhost fits a service behind a local reverse proxy; list the real peer addresses otherwise. For a packaged service, put the directives in a drop-in (systemctl edit nginx) instead of editing the vendor file.

Add one directive at a time on an existing service, restart, and watch journalctl -u app for Permission denied, Read-only file system or Operation not permitted.

Prove it

systemd-analyze verify reported no problem with the unit (the test filters out the complaint that /usr/local/bin/app does not exist in the container). The exposure scores, before and after:

bash
systemd-analyze security --offline=true --no-pager app-plain.service | tail -1
systemd-analyze security --offline=true --no-pager app.service | tail -1
text
→ Overall exposure level for app-plain.service: 9.2 UNSAFE :-{
→ Overall exposure level for app.service: 1.1 OK :-)

What is left, and why each item stays:

bash
systemd-analyze security --offline=true --no-pager app.service | grep "^✗"
text
✗ RootDirectory=/RootImage=                                   Service runs within the host's root directory                                       0.1
✗ PrivateMounts=                                              Service may install system mounts                                                   0.2
✗ RestrictAddressFamilies=~AF_UNIX                            Service may allocate local sockets                                                  0.1
✗ RestrictAddressFamilies=~AF_(INET|INET6)                    Service may allocate Internet sockets                                               0.3
✗ PrivateNetwork=                                             Service has access to the host's network                                            0.5
✗ DeviceAllow=                                                Service has a device ACL with some special devices: char-rtc:r                      0.1
✗ IPAddressDeny=                                              Service defines IP address allow list with only localhost entries                   0.1
✗ UMask=                                                      Files created by service are group-readable by default                              0.1

A web app needs Internet sockets and the host network, so those findings are expected. --threshold turns the score into a CI gate: the command exits non-zero when the exposure is above the number (here 30, meaning 3.0):

bash
systemd-analyze security --offline=true --threshold=30 app-plain.service >/dev/null; echo "exit $?"
systemd-analyze security --offline=true --threshold=30 app.service >/dev/null; echo "exit $?"
text
app-plain.service, threshold 30: exit 1
app.service, threshold 30: exit 0

On a real host, check that the walls hold once the service runs:

bash
systemctl show app -p User -p ProtectSystem -p NoNewPrivileges
cat /proc/$(systemctl show -p MainPID --value app)/status | grep -E "^(NoNewPrivs|CapEff|Seccomp)"
systemd-run --wait --pipe -p ProtectSystem=strict -p DynamicUser=yes touch /etc/probe

What you should see: NoNewPrivs: 1, CapEff: 0000000000000000 and Seccomp: 2 for the main process, and touch: cannot touch '/etc/probe': Read-only file system from the one-off systemd-run. The test script is secure-tests/systemd-service-sandboxing/run.sh.

Mistakes people make

Hardening by copy-paste

A long block copied from a blog can break a service in ways that show up only under load (a JIT, a DNS lookup over netlink, a helper binary). Add directives in groups and test each group.

ProtectSystem=strict without a writable path

The service then fails to write its state or logs. Use StateDirectory=, LogsDirectory=, CacheDirectory= or ReadWritePaths= for the paths it really needs.

Leaving root services unsandboxed

Services that must start as root (to bind port 443, for example) benefit the most. Use AmbientCapabilities=CAP_NET_BIND_SERVICE with a non-root user instead of running as root.

Treating the score as the goal

Getting from 9.2 to 1.1 is useful. Getting from 1.1 to 0.9 by breaking the service's network is not. Stop when the remaining findings are things the service really needs.

Checklist

  • Every custom service runs as a dedicated user or with DynamicUser=yes.
  • ProtectSystem=strict and ProtectHome=yes are set, with explicit writable paths.
  • PrivateTmp=yes and PrivateDevices=yes are set.
  • NoNewPrivileges=yes is set and CapabilityBoundingSet= lists only what is needed.
  • RestrictAddressFamilies= and SystemCallFilter=@system-service are set.
  • systemd-analyze security <unit> scores each custom service, and the scores are recorded.
  • CI runs systemd-analyze security --offline=true --threshold=... on unit files.
  • Vendor units are hardened with drop-ins, not by editing the vendor file.

A sandbox is a promise about the worst day. On that day, you will be glad the service could only reach the one directory it needed.

FND

Learn it on a live range

Linux 2: the kernel, storage, networks and the services a host runs, in Foundation: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on linux hosts

SSH, sudo, firewalls, mount options, updates and the host basics every engineer should get right.

All linux hosts guides