The Kubernetes Security Baseline Nobody Ships: CIS, Pod Security, and the Runtime Layer That Actually Catches Attacks
Sources
- Kubernetes docs — Security concepts overview
- Kubernetes docs — Pod Security Standards (privileged / baseline / restricted)
- Kubernetes docs — Pod Security Admission, namespace labels, exemptions
- Kubernetes docs — Configure the Pod Security Admission Controller (AdmissionConfiguration API)
- Kubernetes docs — AppArmor security profiles (stable since v1.30)
- Kubernetes docs — Restrict a container's syscalls with seccomp (RuntimeDefault)
- Kubernetes docs — Encrypt secrets at rest (provider table, kms v2, secretbox)
- Kubernetes docs — kubelet Configuration (v1beta1) reference: readOnlyPort, protectKernelDefaults, seccompDefault, tlsMinVersion
- CIS Kubernetes Benchmark portal
- kube-bench v0.16.0 — platform support table (CIS 1.12 covers Kubernetes 1.32-1.34)
- kube-bench v0.16.0 — cis-1.12 node checks (kubelet anonymous auth, readOnlyPort, seccomp-default)
- kube-bench v0.16.0 — cis-1.12 policies checks (RBAC, Pod Security Standards section 5.2)
- kube-bench job manifests (hostPID Job with hostPath mounts)
- Trivy v0.75.0 release (2026-10-01) — crypto asset scanning, Echo-patched Python package detection
- Trivy docs — Kubernetes in-cluster scanning (trivy k8s --report summary)
- trivy-operator v0.35.0 (2026-10-05) — continuous in-cluster scanning via Helm
- Tetragon v1.7.1 docs — Kubernetes installation (helm install cilium/tetragon)
- Tetragon v1.7.1 docs — filename-access TracingPolicy (security_file_permission hook)
- Tetragon v1.7.1 metrics reference (tetragon_errors_total, tracingpolicy_loaded, port 2112)
- Falco 0.45.0 + Falco Helm chart (driver.kind=kmod / modern_ebpf)
- Kyverno ValidatingPolicy — disallow-privilege-escalation (CEL, policies.kyverno.io/v1)
- Kyverno ValidatingPolicy — restrict-seccomp-strict (RuntimeDefault or Localhost)
- kube-bench managed-cluster limitation (cannot audit GKE/EKS/AKS control planes)
Every audit season, the same ritual: security asks whether the clusters are "CIS compliant," platform produces a kube-bench run with a pass percentage, and everyone signs the form. Then an attacker gets in through a path the benchmark never touched — a supply-chain image pulled last Tuesday, a debug sidecar with allowPrivilegeEscalation: true that shipped in a Friday hotfix, a kubelet still accepting anonymous requests on a node pool nobody rotated. The CIS Kubernetes Benchmark is a useful checklist of static configuration. It is not a security posture. This guide is the baseline we actually run: three layers — verified configuration (kube-bench), admission-time policy (Pod Security Admission plus Kyverno), and runtime detection (eBPF) — with the manifests we use for each and the reasons each layer exists.
The timing reason this is worth a fresh look: the tooling in this stack just moved. Trivy 0.75.0 shipped October 1, 2026 with cryptographic-asset scanning and detection of Echo-patched Python packages — the class of quietly-rebuilt dependency that fooled SBOM diffing for months. trivy-operator v0.35.0 followed on October 5, making continuous in-cluster scanning effectively hands-off. Tetragon 1.7.1 and Falco 0.45.0 are current stable in the runtime lane, and Kyverno's CEL-based ValidatingPolicy generation — the upstream PSS policies carry a 1.17.0 minimum-version annotation — changed how pod-level policy is written. The baseline below is what those releases make practical.
The Structural Shift: Configuration Checks Don't See Payloads
The uncomfortable truth about a pure CIS posture: every control in the benchmark audits the cluster's own configuration files and flags — permissions on /etc/kubernetes/manifests/kube-apiserver.yaml, whether --anonymous-auth=false is set, whether etcd requires client certificates. Those checks run at a point in time, against components the platform team controls. None of them inspect the thing that actually executes user code: the workload image, its syscalls, its file access. The result is a security boundary that ends at the pod sandbox.
CLUSTER LIFECYCLE WHAT EACH LAYER SEES
─────────────────────────────────────────────────────────────────
kube-apiserver flags & files [kube-bench] static,
kubelet config, etcd TLS (CIS 1.12) point-in-time,
RBAC bindings, PSA labels ──▶ cluster's own configuration-only
configuration
pod admission (CREATE / UPDATE) [PSA + Kyverno] declarative,
securityContext, seccomp, ──▶ the pod spec enforced at the
capabilities, host namespaces API boundary
syscalls, file opens, [eBPF runtime] the actual
process exec, privilege changes ──▶ the workload's payload, live,
behavior in kernel per-event
─────────────────────────────────────────────────────────────────
A cluster can pass 100% of CIS automated checks and let a
curl-piped-into-a-shell container run as uid 0 on day one.That is the structural shift: the Kubernetes security model has three planes — cluster configuration, workload admission, and runtime behavior — and a benchmark only covers the first. The CIS benchmark itself acknowledges this: its Pod Security section (5.2) is almost entirely manual checks — "minimize the admission of privileged containers," "minimize the admission of containers with allowPrivilegeEscalation" — instructions for a human to run kubectl get pods and look. kube-bench marks every one of them Manual, not Automated, because a one-shot audit tool cannot continuously enforce admission. You close that gap with Pod Security Admission and a policy engine, and you close the payload gap with eBPF.
Layer 1 — kube-bench: Run the CIS Scan Correctly (and Know What It Cannot See)
kube-bench v0.16.0 maps Kubernetes versions to CIS benchmark revisions. The current coverage table matters more than the version number: cis-1.12 covers Kubernetes 1.32–1.34, cis-1.11 covers 1.29–1.31 — the table ends at 1.34, so if your minors are newer, check the mapping before assuming coverage rather than trusting auto-detection. The benchmark-to-Kubernetes mapping is explicit in the project docs:
CIS benchmark revision kube-bench config Kubernetes versions
──────────────────────────────────────────────────────────────────
1.10 cis-1.10 1.28
1.11 cis-1.11 1.29 – 1.31
1.12 cis-1.12 1.32 – 1.34
1.7.0-rke2 rke2-cis-1.8/1.9 RKE2 v1.25+
EKS 1.5.0 eks-1.5.0 EKS (worker nodes only)
GKE 1.9.0 gke-1.9.0 GKE (worker nodes only)Run it as a Job. The canonical manifest is a hostPID: true batch Job mounting the host's /etc/kubernetes and /var/lib/kubelet read-only — kube-bench reads kubelet and control-plane config from the node itself, which is why it needs host namespaces and why it should never run with write mounts:
apiVersion: batch/v1
kind: Job
metadata:
name: kube-bench
namespace: security-audit
spec:
template:
metadata:
labels:
app: kube-bench
spec:
hostPID: true
restartPolicy: Never
containers:
- name: kube-bench
image: docker.io/aquasec/kube-bench:v0.16.0
command: ["kube-bench", "run", "--targets", "node", "--include-test-output"]
volumeMounts:
- name: var-lib-kubelet
mountPath: /var/lib/kubelet
readOnly: true
- name: etc-kubernetes
mountPath: /etc/kubernetes
readOnly: true
- name: etc-systemd
mountPath: /etc/systemd
readOnly: true
volumes:
- name: var-lib-kubelet
hostPath:
path: /var/lib/kubelet
- name: etc-kubernetes
hostPath:
path: /etc/kubernetes
- name: etc-systemd
hostPath:
path: /etc/systemdFlags worth knowing, from the tool's own reference: --benchmark cis-1.12 pins the revision instead of auto-detecting from cluster version; --json / --junit with --outputfile feeds a compliance pipeline (or --asff to push findings straight into AWS Security Hub); --exit-code makes failures break CI. Use --include-test-output in any automated run — without it a failed check tells you the flag was wrong but not what it actually found.
The managed-cluster caveat the marketing never prints: kube-bench's own docs state it plainly — it is impossible to inspect the master nodes of GKE, EKS, AKS, or ACK because you do not have access to those nodes. On managed control planes your "CIS compliance" is a worker-node-only score. The upstream job-eks.yaml runs with --targets node,policies,managedservices,controlplane --benchmark eks-1.5.0 — the EKS-tuned profile — but sections 1–3 (control plane node configuration, etcd, control plane configuration) largely fall away because the API server manifest you're supposed to audit is not on your node. Your cloud provider's attestation replaces your evidence. That is a real audit conversation, not a checkbox.
What the kubelet section actually checks (cis-1.12 node profile, from the check definitions themselves): authentication.anonymous.enabled false (4.2.1), authorization-mode not AlwaysAllow (4.2.2), readOnlyPort zero or unset (4.2.4), makeIPTablesUtilChains true (4.2.6), certificate rotation enabled (4.2.10/4.2.11), strong ciphers (4.2.12), and — the one teams miss — 4.2.13: seccomp-default set to true. That last check is the bridge to the next two layers: the benchmark is asking you to make the runtime sandbox strict by default, which is a payload decision, not a configuration-file decision.
Layer 2 — Admission: Pod Security Admission First, Kyverno for the Gaps
Pod Security Admission (PSA) is built into the kube-apiserver since 1.25 stable and replaces the deprecated PSP ecosystem with something much simpler: three policy levels — privileged, baseline, restricted — applied to namespaces via labels, in three modes (enforce, audit, warn) plus a version pin. The levels are defined in the Pod Security Standards: baseline blocks host namespaces, privileged containers, and added capabilities; restricted additionally requires runAsNonRoot, dropped capabilities, and allowPrivilegeEscalation: false.
The production pattern is label-per-namespace, never a cluster-wide default that breaks kube-system:
# New namespace, hardened from birth
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Namespace
metadata:
name: payments
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: v1.37
pod-security.kubernetes.io/audit: restricted
pod-security.kubernetes.io/audit-version: v1.37
pod-security.kubernetes.io/warn: restricted
pod-security.kubernetes.io/warn-version: v1.37
EOF
# Existing namespace: observe first, tighten later
kubectl label ns legacy-team \
pod-security.kubernetes.io/enforce=baseline \
pod-security.kubernetes.io/enforce-version=latest \
pod-security.kubernetes.io/audit=restricted \
pod-security.kubernetes.io/audit-version=latestPinning enforce-version matters more than teams expect: latest silently tightens policy when the control plane upgrades, and a restricted-policy change can block a namespace's next rollout at 2 a.m. Pin to your oldest supported minor, and upgrade the pin deliberately.
PSA exemptions live in the admission controller's AdmissionConfiguration on the API server — there is no Kubernetes API object for them. The config schema, from the official reference:
apiVersion: apiserver.config.k8s.io/v1
kind: AdmissionConfiguration
plugins:
- name: PodSecurity
configuration:
apiVersion: pod-security.admission.config.k8s.io/v1
kind: PodSecurityConfiguration
defaults:
enforce: "privileged"
enforce-version: "latest"
audit: "privileged"
audit-version: "latest"
warn: "privileged"
warn-version: "latest"
exemptions:
usernames: []
runtimeClasses: []
namespaces: []Keep exemptions empty. Every exemption is a permanent hole that outlives the reason it was created — if you need one for a CSI driver or privileged DaemonSet, put it in an explicitly labeled namespace instead, so at least kubectl get ns -L pod-security.kubernetes.io/enforce shows the exception exists.
Where PSA ends and Kyverno begins
PSA validates pod specs against three fixed levels. It cannot express "images must come from our registry," "no latest tags," "seccomp must be RuntimeDefault unless documented," or anything about ConfigMaps, ingress, or RBAC. That gap is why CIS 5.2's manual checks map so cleanly to a policy engine: Kyverno's policy library ships a pod-security-vpol directory that reimplements the PSS restricted set as CEL ValidatingPolicy objects — Kubernetes-native admission (the same ValidatingAdmissionPolicy machinery), no webhook dependency for the core check. The upstream disallow-privilege-escalation policy shows the shape:
apiVersion: policies.kyverno.io/v1
kind: ValidatingPolicy
metadata:
name: disallow-privilege-escalation
annotations:
policies.kyverno.io/title: Disallow Privilege Escalation
policies.kyverno.io/category: Pod Security Standards (Restricted)
policies.kyverno.io/subject: Pod
policies.kyverno.io/minversion: 1.17.0
spec:
validationActions:
- Audit # switch to Deny once the baseline is clean
evaluation:
background:
enabled: true
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["pods"]
variables:
- name: allContainers
expression: >-
object.spec.containers +
object.spec.?initContainers.orValue([]) +
object.spec.?ephemeralContainers.orValue([])
validations:
- expression: >-
variables.allContainers.all(container,
container.?securityContext.allowPrivilegeEscalation.orValue(true) == false)
message: >-
Privilege escalation is disallowed.
All containers must set securityContext.allowPrivilegeEscalation to false.Two production notes on this exact policy. First, validationActions: [Audit] is the correct first state — run the rule for two weeks, count the hits in the policy reports, then flip to Deny per-namespace. Second, the .orValue(true) default in the upstream expression is deliberate: an unset allowPrivilegeEscalation fails the policy — the field's runtime default is true, and Kubernetes forces it to true on any privileged container, so treating "not set" as compliant would be the security hole. That is the PSS restricted semantics, and it is stricter than a hand-rolled "check if set" rule would be. The same pod-security-vpol/restricted directory ships restrict-seccomp-strict (pod and every container must set seccompProfile.type to RuntimeDefault or Localhost — unset is rejected, because unset means the runtime default, which historically was Unconfined), disallow-capabilities-strict, require-run-as-nonroot, and restrict-volume-types. Install the set, shadow PSA restricted where you need custom exceptions, and let PSA handle the rest.
What Kyverno adds over PSA in practice: background scans of existing objects (PSA only sees new admissions), policy reports as a queryable resource, and rules that bind to non-pod kinds — the CIS 5.1 RBAC checks ("minimize wildcard use in Roles and ClusterRoles," "cluster-admin only where required") become enforceable policy instead of a quarterly kubectl review.
Layer 3 — Supply Chain: Continuous Scanning, Not a CI Checkbox
Admission policy governs the pod spec; nothing so far governs the image. The baseline here is trivy-operator running continuously in-cluster, not a Trivy CLI step that gates CI. The difference is blast radius: CI scanning sees images at build time, but clusters accumulate images that were never scanned — mirrored ones, sidecar-injected ones, and the ones deployed before the scanner existed. trivy-operator reconciles VulnerabilityReport and ExposedSecretReport resources against what is actually running:
helm repo add aqua https://aquasecurity.github.io/helm-charts/
helm repo update
helm install trivy-operator aqua/trivy-operator \
--namespace trivy-system \
--create-namespace \
--version 0.37.0
# Reports are Kubernetes objects now — query them like everything else
kubectl get vulnerabilityreports -A
kubectl get exposedsecretreports -A
# The ad-hoc alternative for clusters you don't run continuously:
trivy k8s --report summary
trivy k8s --scanners vuln --report allThe 0.75.0 release (October 1, 2026) is why continuous beats CI-gate right now: it added a full cryptographic-asset model — Trivy now inventories keys, certificates, and signature material inside images and flags weak or expired crypto alongside CVEs — and vulnerability detection for Echo-patched Python packages: dependencies quietly rebuilt by internal mirror tooling that carry different content at the same version. Both are exactly the class of drift that a build-time scan stamps "approved" and a runtime reconciler catches three weeks later when the report regenerates. Pair trivy-operator with admission: Kyverno's verify-images pattern (or a simple "no report = no admission" rule) closes the loop, and the trivy-operator Helm chart's settings cover the config.
Layer 4 — Runtime: eBPF Detection, Where the Baseline Meets Payload
Everything above decides what may enter the cluster. Runtime security observes what payloads do once inside — the layer the CIS benchmark has no position on, and the only layer that catches a compromised pod that arrived with a perfectly clean spec. Two CNCF-ecosystem tools dominate: Tetragon (Cilium's eBPF security agent) and Falco (the original runtime threat detector, 0.45.0 current).
Tetragon installs as a DaemonSet via Helm and hooks kernel-level functions with BPF programs, filtering events in the kernel before they ever reach userspace:
helm repo add cilium https://helm.cilium.io
helm repo update
helm install tetragon cilium/tetragon -n kube-system
kubectl rollout status -n kube-system ds/tetragon -wapiVersion: cilium.io/v1alpha1
kind: TracingPolicy
metadata:
name: passwd-write-detect
spec:
kprobes:
- call: "security_file_permission"
syscall: false
args:
- index: 0
type: "file"
- index: 1
type: "int" # 0x04 MAY_READ, 0x02 MAY_WRITE
selectors:
- matchArgs:
- index: 0
operator: "Equal"
values:
- "/etc/passwd"
- index: 1
operator: "Equal"
values:
- "2" # MAY_WRITE
# add matchActions with action: Sigkill to make this enforcement
# Tetragon filters at the hook: events that don't match never
# leave the kernel, so a noisy cluster doesn't pay userspace cost.That filtering point is the architectural argument for eBPF runtime security over auditd-based approaches: the match happens in a BPF program attached to security_file_permission, so the cost of a non-matching event is a kernel-side check, not a log line shipped, parsed, and discarded. The same policy object can flip from detection to enforcement by adding matchActions — Tetragon can Sigkill the offending process at the syscall — which is why it lands in hardening guides as the enforcement layer, not just telemetry.
Falco takes the broader-ruleset, lower-ops path — default rules ship out of the box, and the Helm chart's driver.kind value is the one decision that matters operationally:
helm repo add falcosecurity https://falcosecurity.github.io/charts
helm repo update
# modern_ebpf: no kernel module build, no headers needed on the node
helm install falco falcosecurity/falco \
--create-namespace \
--namespace falco \
--set driver.kind=modern_ebpf
# kmod: legacy path; needs matching kernel headers on every node
# --set driver.kind=kmodmodern_ebpf removes the kernel-module build step, which historically was the single largest source of Falco install failures on managed node groups with locked-down hosts. Pick kmod only when you have a driver build pipeline you already operate.
The Baseline, As Manifests: What a Hardened Kubelet Looks Like
Layer 1 tells you what to check; here is what passing looks like on the kubelet, where most self-managed clusters bleed points (cis-1.12 section 4.2, mapped to the actual KubeletConfiguration fields the checks read):
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
# 4.2.1 — no anonymous auth
authentication:
anonymous:
enabled: false
x509:
clientCAFile: /etc/kubernetes/pki/ca.crt
# 4.2.2 — never AlwaysAllow
authorization:
mode: Webhook
# 4.2.4 — read-only port off
readOnlyPort: 0
# 4.2.6 — kubelet-managed iptables chains stay on
makeIPTablesUtilChains: true
# 4.2.10/4.2.11 — rotate client and server certs
rotateCertificates: true
featureGates:
RotateKubeletServerCertificate: true
# 4.2.12 — strong TLS floor
tlsMinVersion: VersionTLS12
# 4.2.13 — seccomp strict by default (bridges to Layer 2/3)
seccompDefault: true
# beyond CIS: kernel-tamper guard, blocks sysctl abuse
protectKernelDefaults: trueThree of these deserve comment. seccompDefault: true is the benchmark's own bridge to runtime security — it puts every pod without an explicit profile under RuntimeDefault seccomp instead of Unconfined, which is the single highest-value line in the file for payload containment (and pairs with Kyverno's restrict-seccomp-strict rule, which enforces the same choice at admission). protectKernelDefaults: true makes the kubelet refuse to start on a node whose kernel sysctls are unsafe — cheap insurance on nodes where platform and security teams share ownership. And on managed offerings this entire file may be mostly moot: EKS/GKE/AKS ship hardened kubelet defaults and their CIS attestation is the vendor's, which is precisely why your kube-bench run on those clusters checks worker-node drift, not the control plane.
One control-plane item belongs in the baseline even though kube-bench scores it manual: encryption of secrets at rest. The API server's provider choice has teeth — per the official provider table, aescbc is "not recommended due to CBC's vulnerability to padding oracle attacks," aesgcm requires key rotation every 200,000 writes unless automated, kms v2 is "a good choice if using a third party tool for key management" (stable since v1.29), and secretbox is strong and fast but "uses relatively new encryption technologies." For clusters with a KMS or Vault behind them, kms v2 is the production answer; without one, secretbox with a 32-byte key. Either way the sequence is fixed: new provider first with identity fallback, re-encrypt existing secrets (kubectl get secrets --all-namespaces -o json | kubectl replace -f -), then drop the identity line — the upstream doc spells out the migration order, and skipping the middle step is how clusters end up with secrets that silently fail to decrypt after a key rotation.
Day-2: Metrics, Alerts, and the Honest Limits of Each Layer
What to monitor on the security stack itself. Tetragon's agent exports Prometheus metrics on port 2112 by default in Kubernetes installs (operator metrics on 2113), including tetragon_errors_total (agent health) and tetragon_tracingpolicy_loaded{state=...} — the policy-state gauge matters because a TracingPolicy that fails to load into the kernel reports load_error, not absence: your alert is tracingpolicy_loaded{state="load_error"} > 0, or you will run for months believing a policy protects you while it never attached. trivy-operator reports are Kubernetes resources — monitor report age and scan failure, not just findings (kubectl get vulnerabilityreport -A as a CronJob check is embarrassingly effective). For Falco, monitor driver health per node — a Falco pod that loaded no rules detects nothing and reports readiness anyway in older chart versions.
Runbook notes that bite in production:
- kube-bench output drift: the tool reads node-local config paths from its
cfg/config.yamlper distribution — kubeadm paths, RKE2 paths, Talos paths differ. A cluster provisioned by a nonstandard tool makes kube-bench report false failures (file not found), not passes. Sanity-check one node manually before trending scores. - PSA version-pin upgrades are their own change window: bumping
enforce-versionfrom v1.33 to v1.37 can change what "restricted" means; run the new version inauditmode on the namespace first — the audit annotation surfaces violations without blocking anything. - Kyverno
Audit-to-Denyis per-rule, not global: flip the rules that have flat report counts, leave the noisy ones auditing; a globalDenyon day one is how platform teams get overruled in incident reviews. - eBPF agent on managed control planes: Tetragon and Falco run as DaemonSets on your nodes — on EKS/GKE/AKS they see node workloads, not the managed control plane. That is the same visibility boundary kube-bench hit in Layer 1; two tools, one consistent blind spot, and your cloud provider's control-plane logs are the only coverage there.
- Blast radius of the runtime layer: an eBPF agent with a too-broad policy set costs CPU on every node (kernel-filtered, but the filter itself runs per-event). Start with 5–10 high-signal policies (the /etc/passwd pattern above, exec from writable directories, namespace escapes) — not the full library — and expand on evidence.
Who should skip parts of this: if you run on a managed control plane with no self-hosted nodes, Layer 1 collapses to worker-node checks and vendor attestation — put the effort into Layers 2–4, which is where your actual exposure lives. If your cluster runs a single third-party appliance, even a full four-layer stack is overhead for what is effectively one workload; a hardened namespace and image scanning cover it. And if you cannot staff the Audit-to-Deny review cycle, do not ship Deny rules — an unreviewed deny policy is an outage with a security alibi.
References & Further Reading
- Kubernetes security concepts — the official overview of the three planes this guide layers.
- Pod Security Standards and the Pod Security Admission reference — level definitions, label semantics, exemptions.
- Encrypting secret data at rest — the provider table, kms v2 details, and the migration sequence.
- CIS Kubernetes Benchmark — the current benchmark portal (1.11/1.12 revisions).
- kube-bench platform support — the version-to-benchmark mapping table.
- Trivy and trivy-operator — supply-chain scanning, CLI and in-cluster.
- Tetragon — eBPF-based security observability and runtime enforcement (Cilium project).
- Falco — CNCF runtime security with the default ruleset.
- Kyverno — policy engine; the kyverno/policies library includes the PSS vpol set used above.
- AppArmor in Kubernetes (stable since 1.30, via the
appArmorProfilefield) and seccomp — the per-pod kernel confinement this baseline leans on.