Nodes and clusters

Pod Security Admission levels in practice, the secure way

kubectl said "deployment.apps/app created", the pipeline went green, and the Deployment has sat at zero ready replicas for an hour. Pod Security Admission did its job; it just told the ReplicaSet instead of you.

The short answer

Label every application namespace with enforce, warn and audit at restricted, each pinned to a Kubernetes minor version. Keep privileged namespaces few and named. Do not exempt users or RuntimeClasses. Test label changes with a server-side dry run, write workloads that pass restricted, and bump the pinned version on purpose after each upgrade.

Updated Houssam Hammoudi, CTOTested with Kubernetes 1.34.0 (kind); Talos patch checked against the Talos 1.14 docs

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Pod Security Admission (PSA) checks pods against three profiles: privileged (no checks), baseline (blocks known escalations such as host namespaces and privileged containers) and restricted (also requires non-root, no privilege escalation, dropped capabilities and a seccomp profile). A namespace chooses a profile per mode with labels: enforce rejects, warn prints a warning to the client, audit adds an annotation to the audit event.

The traps are in how the modes behave:

Enforce skips workload objects. PSA enforces on pods, not on Deployments, Jobs or StatefulSets. A Deployment that violates the profile is accepted. Its ReplicaSet then fails to create pods, and the only trace is a FailedCreate event on the ReplicaSet. kubectl apply shows a warning only if warn is at least as strict as enforce. A namespace with an enforce label and no warn label gets warn at the enforce level automatically, but a warn label set lower, or an enforce level that comes only from the cluster default while the default warn is lower, leaves kubectl silent. A warning never fails kubectl apply either, so the pipeline stays green.

latest moves. A mode without a -version label uses the cluster default version, which is latest unless the admission configuration pins it. latest means the newest definition of the profile. When a Kubernetes upgrade tightens a profile, pods that were fine yesterday are rejected the next time they are created: a rollout, a node drain, a scale-up.

Exemptions are wide. An exempt RuntimeClass exempts every pod that names it, whoever creates it. An exempt user exempts only pods that user creates directly, not pods created by controllers, so it does not do what people expect and tempts them to exempt the controllers next.

Everything else is baseline. Talos enforces baseline cluster-wide with restricted only in warn and audit. That blocks the worst, but lets containers run as root with every default capability.

What the docs say

However, enforce mode is not applied to workload resources, only to the resulting pod objects.

Source: Kubernetes docs, Pod Security Admission

RuntimeClassNames: pods and workload resources specifying an exempt runtime class name are ignored.

Source: Kubernetes docs, Pod Security Admission

It is helpful to apply the --dry-run flag when initially evaluating security profile changes for namespaces.

Source: Kubernetes docs, Enforce Pod Security Standards with Namespace Labels

The docs explain each mode, but not what the combination looks like from a pipeline: a green kubectl apply and no pods. They also do not say that a missing warn label follows a stricter enforce label; that rule is in the PolicyToEvaluate function of the Pod Security Admission source. Talos generates enforce-version: latest in its cluster default, so every Kubernetes upgrade can change what is enforced without a config change.

The secure configuration

1. Application namespaces: restricted in all three modes, pinned.

yaml
# k8s-namespace.yaml
apiVersion: v1
kind: Namespace
metadata:
  name: app
  labels:
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/enforce-version: v1.37   # bump on purpose after an upgrade
    pod-security.kubernetes.io/warn: restricted          # kubectl apply shows violations
    pod-security.kubernetes.io/warn-version: v1.37
    pod-security.kubernetes.io/audit: restricted         # audit events carry the violation
    pod-security.kubernetes.io/audit-version: v1.37

2. A workload that passes restricted. The comments mark the fields the profile requires; the rest is cheap hardening:

yaml
# k8s-deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
  namespace: app
spec:
  replicas: 2
  selector:
    matchLabels:
      app: web
  template:
    metadata:
      labels:
        app: web
    spec:
      automountServiceAccountToken: false   # not a PSA rule; see the ServiceAccount page
      securityContext:
        runAsNonRoot: true                  # restricted: must not run as UID 0
        runAsUser: 10001
        runAsGroup: 10001
        fsGroup: 10001
        seccompProfile:
          type: RuntimeDefault              # restricted: RuntimeDefault or Localhost
      containers:
        - name: web
          image: ghcr.io/example/web@sha256:0000000000000000000000000000000000000000000000000000000000000000
          ports:
            - containerPort: 8080           # unprivileged port; no NET_BIND_SERVICE needed
          securityContext:
            allowPrivilegeEscalation: false # restricted: must be false
            readOnlyRootFilesystem: true    # not required by PSA, but cheap
            capabilities:
              drop: ["ALL"]                 # restricted: drop ALL (only NET_BIND_SERVICE may be added)
          resources:
            requests:
              cpu: 50m
              memory: 64Mi
            limits:
              memory: 128Mi
          volumeMounts:
            - name: tmp
              mountPath: /tmp
      volumes:
        - name: tmp
          emptyDir:
            sizeLimit: 64Mi

3. The cluster default, pinned, with no RuntimeClass or user exemptions. On Talos 1.14, talosctl gen config already writes a PodSecurity document whose exemptions are only the kube-system namespace, with empty runtimeClasses and usernames. The patch below merges into that document. Lists are appended when patches merge, so do not repeat the exemptions: a second kube-system entry is a duplicate, and kube-apiserver rejects a PodSecurity configuration with duplicate exemptions.

yaml
# talos-psa-defaults.yaml
# talosctl gen config ... --config-patch-control-plane @talos-psa-defaults.yaml
apiVersion: v1alpha1
kind: KubeAdmissionControlConfig
name: PodSecurity
configuration:
  apiVersion: pod-security.admission.config.k8s.io/v1
  kind: PodSecurityConfiguration
  defaults:
    enforce: baseline           # namespaces without labels get at least baseline
    enforce-version: v1.37
    warn: restricted
    warn-version: v1.37
    audit: restricted
    audit-version: v1.37
  # No exemptions block: keep the generated one (kube-system only, no
  # runtimeClasses, no usernames). Check it with: talosctl get
  # admissioncontrolconfigs -o yaml (or read the rendered machine config).

Namespaces that must run privileged agents (CNI, CSI, node monitoring) get enforce: privileged explicitly, are listed in a review document, and allow no user workloads.

4. Change labels with a dry run first.

bash
kubectl label --dry-run=server --overwrite ns app \
  pod-security.kubernetes.io/enforce=restricted \
  pod-security.kubernetes.io/enforce-version=v1.37

Prove it

Run on a lab cluster (Kubernetes 1.34) with the manifests above; the Deployment's placeholder image was replaced by busybox:1.37 serving HTTP on 8080, with every security setting as published. The Talos defaults patch was not applied (the lab is not Talos). Real output:

1. The dry run lists what would break, and changes nothing:

bash
kubectl label --dry-run=server --overwrite ns --all pod-security.kubernetes.io/enforce=restricted
text
namespace/app not labeled (server dry run)
namespace/cilium-secrets labeled (server dry run)
namespace/default labeled (server dry run)
Warning: existing pods in namespace "kube-system" violate the new PodSecurity enforce level "restricted:latest"
Warning: coredns-66bc5c9577-snqbj (and 1 other pod): runAsNonRoot != true, seccompProfile
Warning: etcd-lab-control-plane (and 3 other pods): host namespaces, hostPort, probe or lifecycle host, allowPrivilegeEscalation != false, unrestricted capabilities, restricted volume types, runAsNonRoot != true

app already enforces restricted, so it is "not labeled". The system pods fail as expected: that is why kube-system stays exempt.

2. A bad Deployment warns at apply time:

bash
kubectl -n app create deployment bad --image=busybox:1.37 -- sleep 3600
text
Warning: would violate PodSecurity "restricted:v1.37": allowPrivilegeEscalation != false (container "busybox" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "busybox" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "busybox" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "busybox" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")
deployment.apps/bad created

3. And its pods never appear:

bash
kubectl -n app get deployment bad
kubectl -n app get events --field-selector reason=FailedCreate
text
NAME   READY   UP-TO-DATE   AVAILABLE   AGE
bad    0/1     0            0           8s
FailedCreate   Error creating: pods "bad-854ccd948f-txfzx" is forbidden: violates PodSecurity "restricted:v1.37": allowPrivilegeEscalation != false ...

Delete it with kubectl -n app delete deployment bad.

4. The good Deployment runs:

bash
kubectl apply -f k8s-deployment.yaml
kubectl -n app rollout status deployment/web
text
deployment.apps/web created
Waiting for deployment "web" rollout to finish: 0 of 2 updated replicas are available...
Waiting for deployment "web" rollout to finish: 1 of 2 updated replicas are available...
deployment "web" successfully rolled out

No warning. With the placeholder image left in, the pods are admitted and then sit in ImagePullBackOff: admission passed, the image did not exist.

Mistakes people make

Warn set lower than enforce

The Deployment is accepted, the pods are not, and the pipeline reports success. Set warn explicitly to the same level and version as enforce instead of relying on defaults, and check ReplicaSet events in the deploy job, because a warning alone never fails it.

Leaving the version at latest

A Kubernetes upgrade can change what restricted means. With a pinned version, the upgrade changes nothing until you bump the label, after a dry run.

Exempting a RuntimeClass for the sandbox

It seems natural to exempt gvisor because sandboxed pods are "safe". Any pod that sets runtimeClassName: gvisor then skips every check, including hostPath and host namespaces, which gVisor does not make safe. Use a labeled namespace instead.

Exempting the user who deploys

User exemptions cover only pods that user creates directly. Deployments are turned into pods by the ReplicaSet controller, so the exemption does nothing, and the next step, exempting the controller, disables PSA for everyone.

Giving teams namespace update rights

Whoever can update a Namespace can change its PSA labels to privileged. The built-in admin and edit roles cannot update Namespaces; keep it that way, and review any role that grants update or patch on namespaces.

Checklist

  • Every application namespace has enforce, warn and audit at restricted.
  • Every mode has a -version label pinned to a minor version.
  • The cluster default enforces at least baseline, with a pinned version.
  • Exemptions list only kube-system; no RuntimeClasses and no usernames.
  • Privileged namespaces are few, listed and reviewed.
  • Label changes are tested with --dry-run=server first.
  • Deploy jobs fail on FailedCreate ReplicaSet events.
  • Workloads set runAsNonRoot, seccompProfile: RuntimeDefault, allowPrivilegeEscalation: false and drop ALL capabilities.
  • Only platform admins can update or patch Namespace objects.

Pod Security Admission never lies; it just talks to the ReplicaSet. Turn on `warn`, and it will talk to you too.

H2-CSPE

Learn it on a live range

Immutable OS and cluster hardening, in Secure Platform Engineering: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on nodes and clusters

Talos Linux, Kubernetes API hardening, service account tokens, RBAC and the cloud underneath.

All nodes and clusters guides