Linux hosts

Container hosts: rootless containers, the secure way

You added a developer to the docker group so they would stop asking for sudo. Congratulations: they no longer need sudo, because the docker group is root with extra steps and a friendlier logo.

The short answer

Run containers rootless: each user runs Podman (or rootless Docker) with a unique subordinate UID range in /etc/subuid and /etc/subgid. Root inside the container maps to the user outside it, so a container escape lands as an ordinary user. Do not put people in the docker group; that group can become root on the host.

Updated Houssam Hammoudi, CTOTested with Podman 4.3.1 (Debian 12)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

A classic Docker setup runs one daemon as root. Anyone who can talk to its socket can ask it to start a container with the host's root filesystem mounted, and then read or change any file on the host. That is why membership in the docker group is equal to root.

The second problem is what "root in a container" means. In a rootful container, UID 0 inside is UID 0 on the host, only fenced in by namespaces, capabilities and seccomp. A kernel bug or a bad mount turns a container escape into host root.

Rootless containers change both. There is no root daemon, and root inside the container is mapped to the ordinary user who started it. An escape lands as that user.

What the docs say

First of all, only trusted users should be allowed to control your Docker daemon.

Source: Docker docs, Docker Engine security

Rootless Podman is not, and will never be, root; it's not a setuid binary, and gains no privileges when it runs.

Source: Podman rootless tutorial

If there is overlap, there is a potential for a user to use another user's namespace and they could corrupt it.

Source: Podman rootless tutorial, /etc/subuid and /etc/subgid configuration

The Docker sentence is polite about it. In practice "control your Docker daemon" means root on the host. The Podman tutorial is clear that ranges must not overlap; most hosts never check.

The secure configuration

bash
# Packages (Debian and Ubuntu; on Fedora and RHEL: dnf install podman).
apt-get install -y podman uidmap slirp4netns fuse-overlayfs

# Each user gets a unique, non-overlapping range of 65536 IDs.
# useradd does this automatically; for existing users:
usermod --add-subuids 100000-165535 --add-subgids 100000-165535 alice
grep alice /etc/subuid /etc/subgid

# Nobody runs containers through a root daemon socket.
getent group docker        # should be empty, or the group should not exist

# Services in rootless containers keep running after logout.
loginctl enable-linger alice

As the user:

bash
podman info --format '{{.Host.Security.Rootless}}'   # true
podman run --rm docker.io/library/alpine:3.22 id
podman unshare cat /proc/self/uid_map                 # the ID mapping

For each container, keep the rest of the defenses too: run the app as a non-root user inside the image, drop capabilities (--cap-drop=all), and use --read-only where you can. Rootless limits the damage of an escape; it does not replace these.

If you must keep Docker, use its rootless mode, or enable userns-remap so container root maps to an unprivileged host range.

Prove it

The subordinate ranges useradd created, and Podman's own view:

bash
grep alice /etc/subuid /etc/subgid
podman info --format "rootless={{.Host.Security.Rootless}} graphDriver={{.Store.GraphDriverName}}"
text
/etc/subuid:alice:100000:65536
/etc/subgid:alice:100000:65536
rootless=true graphDriver=vfs

Inside the container, the process is root:

bash
podman run --rm -q docker.io/library/alpine:3.22 id
text
uid=0(root) gid=0(root) groups=0(root),1(bin),2(daemon),3(sys),4(adm),6(disk),10(wheel),11(floppy),20(dialout),26(tape),27(video)

The map explains why that is harmless: container UID 0 is host UID 1000 (alice), and container UIDs 1 and up are alice's range starting at 100000:

bash
podman unshare cat /proc/self/uid_map
text
         0       1000          1
         1     100000      65536

On the host, the "root" process of a running container belongs to alice:

bash
ps -eo user,pid,args | awk 'NR==1 || /sleep 300/ && !/awk/'
text
USER         PID COMMAND
alice       4163 sleep 300

Root inside the container cannot read a root-only host file mounted into it:

bash
podman run --rm -q -v /srv:/host:ro docker.io/library/alpine:3.22 sh -c "id -u; cat /host/host-root-only"
text
0
cat: can't open '/host/host-root-only': Permission denied

The storage driver was vfs because the test ran nested inside another container; on a normal host Podman uses overlay. The test script is secure-tests/container-hosts-rootless-containers/run.sh.

Mistakes people make

The docker group as a convenience

It is a root shell for anyone in it. Remove people from it, and give them rootless Podman or rootless Docker instead.

Overlapping subuid ranges

Two users with overlapping ranges can reach each other's container files. Check /etc/subuid and /etc/subgid for overlaps when you add users by hand.

Thinking rootless means safe

A rootless container can still read everything the user can read, including their SSH keys and tokens. Run workloads under dedicated service users, not under a person's account.

--privileged "because it works"

--privileged in rootless mode gives the container every capability inside its user namespace and access to the user's devices. It is less dangerous than rootful --privileged, and still far more than an app needs.

Disabling user namespaces for hardening, then enabling root containers

Some hardening guides set user.max_user_namespaces = 0. On a container host that forces everything back to rootful containers. Pick one on purpose.

Checklist

  • Containers are started by ordinary users with rootless Podman or rootless Docker.
  • Every user who runs containers has a unique range in /etc/subuid and /etc/subgid.
  • No ranges overlap.
  • The docker group is empty or does not exist.
  • podman info reports rootless: true for each service account.
  • Workloads run under dedicated service users, with loginctl enable-linger where needed.
  • Containers still drop capabilities and run as a non-root user in the image.
  • No container runs with --privileged unless there is a written reason.

Rootless containers do not make escapes impossible. They make an escape land in a small room instead of the control tower.

FND

Learn it on a live range

Containers and Kubernetes, in Foundation: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on linux hosts

SSH, sudo, firewalls, mount options, updates and the host basics every engineer should get right.

All linux hosts guides