Runtime detection and observability

Kubescape CIS, NSA and MITRE scans, the secure way

The cluster passed its audit last spring. Since then, forty pull requests added a privileged pod, a hostPath mount of the root disk and a container running as root, each one "just for now". Audits look backward; a scanner in CI looks at the next change.

The short answer

Run kubescape scan framework nsa,mitre on every manifest change in CI and fail the build with --compliance-threshold. Fix what it finds in the manifests: no privileged containers or hostPath, non-root, no privilege escalation, read-only root, dropped capabilities, limits, and a network policy. Run the CIS framework against the live cluster with the Kubescape Operator for node checks.

Updated Houssam Hammoudi, CTOTested with Kubescape 4.0.14 (checksum-verified release binary)

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

Kubernetes accepts almost any workload you give it. A Deployment can ask for a privileged container, mount the node's root filesystem, run as root, and skip resource limits. Each of those is one line of YAML and often the quickest way past an error.

Point-in-time audits find these months later. By then the risky setting is load-bearing, and removing it breaks something.

Frameworks such as the NSA-CISA Kubernetes Hardening Guide, the MITRE ATT&CK matrix for Kubernetes and the CIS Kubernetes Benchmark describe what "secure" means. Kubescape turns them into automated checks. The trick is to run the right checks in the right place: workload checks on manifests before merge, node and control-plane checks on the running cluster.

What the docs say

Kubescape can be used to scan local YAML/JSON files before they are deployed to Kubernetes.

Source: Kubescape docs, Scanning your environment

Some controls require validating a configuration which has to run on the nodes of a Kubernetes cluster; for example, to check if the Kubelet enforces client TLS authentication. To enable the host scanner, install the Kubescape Operator in your cluster.

Source: Kubescape docs, Scanning your environment, The host scanner

The docs list the frameworks and the flags. They do not say which framework fits which stage. NSA and MITRE work well on manifests. Most CIS checks are about kubelet and control-plane settings, which exist only in a running cluster.

The secure configuration

The workload before:

yaml
# before.yaml: a Deployment the way many first drafts look.
apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
  namespace: team-a
spec:
  replicas: 2
  selector:
    matchLabels: {app: web}
  template:
    metadata:
      labels: {app: web}
    spec:
      containers:
        - name: web
          image: nginx:latest
          ports: [{containerPort: 80}]
          securityContext:
            privileged: true
          volumeMounts:
            - {name: host, mountPath: /host}
      volumes:
        - name: host
          hostPath: {path: /}

The same workload after:

yaml
# after.yaml: the same workload, hardened.
apiVersion: apps/v1
kind: Deployment
metadata:
  name: web
  namespace: team-a
spec:
  replicas: 2
  selector:
    matchLabels: {app: web}
  template:
    metadata:
      labels: {app: web}
    spec:
      automountServiceAccountToken: false
      securityContext:
        runAsNonRoot: true
        runAsUser: 101
        runAsGroup: 101
        seccompProfile: {type: RuntimeDefault}
      containers:
        - name: web
          # pinned by digest: the tag can move, the digest cannot
          image: nginxinc/nginx-unprivileged:1.29-alpine@sha256:0c79d56aee561a1d81c63f00eee5fb5fe29279560cdc55e91425133104c7fbe6
          ports: [{containerPort: 8080}]
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities: {drop: [ALL]}
          resources:
            requests: {cpu: 50m, memory: 64Mi}
            limits: {cpu: 500m, memory: 128Mi}
          livenessProbe: {httpGet: {path: /, port: 8080}}
          readinessProbe: {httpGet: {path: /, port: 8080}}
          volumeMounts:
            - {name: tmp, mountPath: /tmp}
      volumes:
        - name: tmp
          emptyDir: {}
---
# Default-deny for the namespace, then allow only what web needs.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: web
  namespace: team-a
spec:
  podSelector:
    matchLabels: {app: web}
  policyTypes: [Ingress, Egress]
  ingress:
    - from:
        - namespaceSelector:
            matchLabels: {kubernetes.io/metadata.name: ingress}
      ports: [{port: 8080}]
  egress: []

The CI step. It fails the build when the combined score is below 90%:

bash
# Verify the release checksum once when you install the binary in your CI image.
kubescape scan framework nsa,mitre deploy/*.yaml \
  --compliance-threshold 90 \
  --format junit --output kubescape.xml

On the running cluster, schedule the CIS scan with the Kubescape Operator (installed with Helm in its own namespace), which adds the host scanner:

bash
kubescape scan framework cis-v1.12.0 --verbose

Prove it

NSA framework on before.yaml:

bash
kubescape scan framework nsa before.yaml --logger error --format pretty-printer
text
│ Severity │ Control name                         │ Failed resources │ All Resources │ Compliance score │
├──────────┼──────────────────────────────────────┼──────────────────┼───────────────┼──────────────────┤
│   High   │ Privileged container                 │         1        │       1       │        0%        │
│   High   │ Ensure CPU limits are set            │         1        │       1       │        0%        │
│   High   │ Ensure memory limits are set         │         1        │       1       │        0%        │
│  Medium  │ Non-root containers                  │         1        │       1       │        0%        │
│  Medium  │ Allow privilege escalation           │         1        │       1       │        0%        │
│  Medium  │ Ingress and Egress blocked           │         1        │       1       │        0%        │
│  Medium  │ Automatic mapping of service account │         1        │       1       │        0%        │
│  Medium  │ Linux hardening                      │         1        │       1       │        0%        │
│    Low   │ Immutable container filesystem       │         1        │       1       │        0%        │
├──────────┼──────────────────────────────────────┼──────────────────┼───────────────┼──────────────────┤
│          │           Resource Summary           │         1        │       1       │      55.00%      │
╰──────────┴──────────────────────────────────────┴──────────────────┴───────────────┴──────────────────╯

MITRE framework on before.yaml. It adds the two findings that matter most here, the hostPath mount of /:

text
│ Severity │ Control name            │ Failed resources │ All Resources │ Compliance score │
├──────────┼─────────────────────────┼──────────────────┼───────────────┼──────────────────┤
│   High   │ Writable hostPath mount │         1        │       1       │        0%        │
│   High   │ HostPath mount          │         1        │       1       │        0%        │
│   High   │ Privileged container    │         1        │       1       │        0%        │
├──────────┼─────────────────────────┼──────────────────┼───────────────┼──────────────────┤
│          │     Resource Summary    │         1        │       1       │      82.35%      │
╰──────────┴─────────────────────────┴──────────────────┴───────────────┴──────────────────╯

Both frameworks on after.yaml: every control passes (the full tables list each control at 100%; the summary lines):

text
nsa:   │          │                    Resource Summary                   │         0        │       2       │      100.00%     │
mitre: │          │                    Resource Summary                   │         0        │       1       │      100.00%     │

The CI gate, with --compliance-threshold 90:

text
before.yaml: exit 1
after.yaml: exit 0

The CIS scan of a live cluster was not run for this page. What you should see: a table of CIS controls grouped by section (control plane, etcd, kubelet, policies), with node checks such as kubelet authentication filled in only when the Kubescape Operator's host scanner is installed; without it, those controls are skipped. The test script is secure-tests/kubescape-cis-nsa-mitre-scans/run.sh.

Mistakes people make

Scanning only the live cluster

By the time a scan of the cluster finds a privileged pod, it is running. Scan the manifests in CI so the finding is a review comment, not an incident.

Running CIS on manifests

Most CIS Kubernetes Benchmark items check kubelet, API server and etcd settings. A manifest scan cannot see them. Use NSA and MITRE for manifests, CIS for clusters.

A threshold of zero

A gate that never fails is decoration. Start with the current score, raise it as you fix findings, and never lower it without a written exception.

Exceptions without an owner

Kubescape supports exceptions. Each one needs a reason, an owner and an expiry date, or the list grows until the scan means nothing.

Pointing the CLI at production from a laptop

A cluster scan uses your current kubeconfig. Run scheduled scans in the cluster with the operator, or from CI with a read-only service account.

Checklist

  • Every change to Kubernetes manifests runs kubescape scan framework nsa,mitre in CI.
  • The build fails below a compliance threshold that only goes up.
  • No workload runs privileged or mounts hostPath without a written exception.
  • Workloads run as non-root with allowPrivilegeEscalation: false and dropped capabilities.
  • Every container has CPU and memory limits.
  • Every namespace has a default-deny network policy.
  • The CIS framework runs on the live cluster with the operator's host scanner.
  • Every exception has an owner and an expiry date.

A benchmark is only useful if it runs more often than your auditors do. Put it where the YAML changes.

H2-CTDE

Learn it on a live range

Detection engineering, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on runtime detection and observability

Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.

All runtime detection and observability guides