Kubernetes Node Drains at Scale: PDB Math, 429 Storms, and the EvictionRequest API That Fixes Both
Sources
- Kubernetes docs — API-initiated Eviction (Eviction API, 429 TooManyRequests, stuck evictions)
- Kubernetes docs — Specifying Disruptions (PodDisruptionBudget status fields, percentage rounding)
- KEP-4563: EvictionRequest API — sig-node, alpha in 1.37, feature gate EvictionRequestAPI
- Kubernetes API reference — EvictionRequest (lifecycle.k8s.io/v1alpha1): intent, requester, target, TargetEvicted/Failed conditions
- Kubernetes API reference — Eviction (policy/v1): subresource endpoint, deleteOptions, Retry-After semantics
- kubectl reference — drain flags: --disable-eviction, --delete-emptydir-data, --grace-period, --timeout, --pod-selector, --ignore-errors
- Kubernetes issue #114877 — maxSurge for node draining (pacing drains by adding capacity)
- Kubernetes issue #75957 — cannot drain node with pod matched by more than one PDB
- Kubernetes issue #91492 — preemption can violate PDBs
- Kubernetes issue #106286 — confusing use of TooManyRequests for eviction
- Karpenter NodePool API — disruption budgets (default 10%), terminationGracePeriod bypassing PDBs
- Cluster Autoscaler FAQ — eviction retries up to 2 min per node, PDB respect in scale-down
- EKS managed node group update behavior — updateConfig.maxUnavailable, 10-node parallel cap
- kube-state-metrics — poddisruptionbudget metrics (disruptions allowed, desired healthy)
- Kubernetes docs — Graceful node shutdown (kubelet shutdownGracePeriod, two-phase, systemd)
Nobody writes a blog post about the day the drain worked. The stories that reach your incident channel are the other kind: an upgrade wave that was supposed to finish in two hours and is still wedged at node 412 of 640 eight hours later, with a postgres pod that refuses to move and a PDB that says DisruptionsAllowed: 0 with total sincerity. This guide is about why that happens in 1,000+ node clusters specifically — the arithmetic, the client behaviors, and the controller mechanics that are invisible at 40 nodes and load-bearing at 1,000 — and what the new EvictionRequest API (KEP-4563, alpha in Kubernetes 1.37) actually changes.
First, the part everyone gets wrong: a node drain is not a kubectl feature — it is a loop over the Eviction API, and the Eviction API is a request-response endpoint over a PodDisruptionBudget admission check. kubectl drain is a serial client walking that endpoint one pod at a time, honoring PDBs, honoring graceful termination, and never — by design — telling the API server "I need this whole node empty by 02:00." The official eviction docs are candid about the consequence: when a pod's replacement never becomes Ready, the eviction keeps returning 429 Too Many Requests "until you intervene." At 40 nodes, intervention is an SRE with a terminal. At 1,000 nodes, intervention is a job title.
The Structural Shift: From "Delete the Pods" to "Coordinated, Budgeted Disruption"
The original Kubernetes answer to "empty this node" was kubectl delete pod --all-namespaces --field-selector spec.nodeName=... — destructive, PDB-blind, and fast. Everything since has been a long retreat from that position, adding layers of coordination between your intent ("this node goes away at 02:00") and the pod's fate. Count the layers that a modern drain passes through:
your intent: "node-7 dies tonight"
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ drain orchestrator (kubectl / Karpenter / CA / EKS updater) │
│ decides WHICH node, WHEN, and how to pace parallel drains │
└─────────────────────────────────────────────────────────────────┘
│ POST /api/v1/namespaces/ns/pods/NAME/eviction
▼
┌─────────────────────────────────────────────────────────────────┐
│ kube-apiserver: Eviction API (policy/v1 subresource) │
│ • admission: find PDBs matching the pod │
│ • if budget would be violated → 429 TooManyRequests │
│ • success → pod deletionTimestamp set, graceful shutdown clock│
└─────────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────────┐
│ PodDisruptionBudget controller (kube-controller-manager) │
│ recomputes DisruptionsAllowed every sync from live pod health │
└─────────────────────────────────────────────────────────────────┘
│ terminationGracePeriodSeconds counts down
▼
┌─────────────────────────────────────────────────────────────────┐
│ kubelet on the old node: preStop hooks → SIGTERM → SIGKILL │
│ (Graceful Node Shutdown adds two-phase ordering on reboots) │
└─────────────────────────────────────────────────────────────────┘
│ endpoints/EndpointSlice controller removes the IP
▼
traffic drains; scheduler places the replacement elsewhereEvery arrow in that diagram is a place where a 1,000-node drain dies. Not the arrows themselves — the coordination gaps between them, which is exactly the framing the KEP-4563 authors use: today's tooling drives API-initiated eviction "in an application agnostic way," and PDBs — the application's only standard defense — were never designed to express "evict this pod when you can, not now or never."
Layer 1: PDB math is status math, not spec math
The budget that stalls your drain is computed from live cluster state, not from the YAML you wrote. Per the PDB documentation, the status fields are:
currentHealthy 3 ← number of pods reporting Ready (live)
desiredHealthy 2 ← spec-derived: expectedPods - allowed disruptions
expectedPods 3 ← total pods matched by .spec.selector
disruptionsAllowed 1 ← currentHealthy - desiredHealthy
observedGeneration 1The formula from the same docs: disruptionsAllowed = currentHealthy - desiredHealthy. That is the entire gate. Note what is not in it: your deadline, your wave plan, your upgrade window. The PDB does not know a drain is happening. It only knows how many pods are Ready right now — which is why a drain that would be trivially fine at steady state becomes wedged when the replacement pods from the previous wave haven't passed their probes yet.
Two edge cases from the same page bite harder at scale. Percentage values: minAvailable: "50%" with 7 pods rounds up to 4 — so small workloads with percentage PDBs lose more budget than you'd guess. And a maxUnavailable: 0% with zero expected pods produces DisruptionsAllowed: 0 — the classic "PDB covers zero pods and blocks everything" trap. The docs' own example of this: a single-instance app where the PDB admits no disruption at all; their stated solution is to use minAvailable: "90%" instead, or drop the PDB entirely.
Layer 2: kubectl drain is a serial client, and serial does not scale
kubectl drain walks the node's pods one eviction at a time. It is a one-node tool with honest defaults, and the generated reference shows its full flag surface:
--disable-eviction=false Force drain to use delete, bypassing PDB checks
--delete-emptydir-data=false Continue even if pods use emptyDir
--force=false Continue even if pods have no controller
--grace-period=-1 Use the pod's own terminationGracePeriodSeconds
--ignore-daemonsets=false Ignore DaemonSet-managed pods
--pod-selector='' Filter pods by label selector
--skip-wait-for-delete-timeout=0 Skip waiting for pods already deleting >N s
--timeout=0s Give up after this long; zero means infiniteRead --timeout 0s — zero means infinite again, because that is the default posture of the tool most teams reach for during an incident: it will sit under a blocked eviction forever, and the docs warn explicitly that evictions can be stuck "if the last evicted Pod had a long termination grace period" or if new pods never become Ready. None of these flags let you say "pace this across the fleet" or "abort this wave at 04:00." Those are orchestrator concerns, and the orchestrators each answer them differently — EKS caps parallel node drains at 10, Karpenter budgets them by percentage, Cluster Autoscaler retries for 2 minutes and then moves on.
The Numbers That Break at 1,000 Nodes
At 40 nodes, drain failure is a rounding error. At 1,000, every per-pod constant multiplies into a wall-clock problem. The three that matter:
| Constant | Value | Where it comes from | At 1,000 nodes |
|---|---|---|---|
| Serial eviction per node | 1 pod at a time | kubectl drain client loop | Pods per node × (PDB wait + grace period) stacks up per node, serially |
| PDB budget refresh cadence | disruption controller sync | kube-controller-manager | A blocked PDB stays blocked until replacements become Ready — no back-pressure to the drain |
| Eviction API response on violation | 429 TooManyRequests | API-initiated eviction docs | Retry storms across parallel drainers multiply API server load |
Now multiply it out with a worked example. This is the arithmetic of a stalled upgrade wave — run the same numbers for your own fleet:
# Assumptions (verify against your own cluster):
# 1000 nodes, 10 pods/node worth evicting, PDB allows 1 disruption per workload
# kubectl drain --timeout 15m per node (no infinite wait), serial within node
# replacement pods need 120s to become Ready (probe delays)
#
# Steady state, everything healthy:
# per pod: eviction accepted → grace period 30s → replacement Ready at +120s
# per node: 10 pods, but PDB allows 1 at a time → ~10 × 150s = 25 min/node
# PDB refill: next disruption allowed once replacement becomes Ready
#
# The wedge: any pod whose replacement never becomes Ready
# → eviction returns 429 forever (docs: "until you intervene")
# → kubectl drain --timeout expires → node stays cordoned, still occupied
# → cordoned nodes never rejoin rotation → capacity shrinks silently
#
# Second-order: parallel drainers all polling the same starved PDBs
# → 429 storm on the API server (also indistinguishable from rate limiting)
# → your automation sees "429" and cannot tell budget-block from throttleThe last line is the operational kicker the KEP authors called out directly: issue #106286 — "Confusing use of TooManyRequests error for eviction" — is one of the standing PDB issues the KEP lists. Your runbook cannot distinguish "the application is unhealthy" from "you're hitting API priority and fairness limits" because the API uses the same status code for both.
The multi-PDB deadlock
The other classic wedge: one pod selected by two PDBs. Per issue #75957 ("Cannot drain node with pod with more than one Pod Disruption Budget"), the eviction admission path checks each matching PDB and fails if any single budget would be violated — so overlapping selectors (say, a shared team=payments budget plus a per-app budget) can produce a pod that can never be evicted while any one of its budgets is tight, even when the other has budget to spare. The KEP's standing-issues list treats this as a defect class to design out, not a misconfiguration.
Where orchestrators already pace drains — and where they cap you
Nobody sane drains 1,000 nodes with kubectl. The real fleets use orchestrators, and their pacing is the actual "drain policy" of your cluster, whether you've read it or not:
| Orchestrator | Drain/pacing control | Default / cap | Gotcha |
|---|---|---|---|
| EKS managed node groups | updateConfig.maxUnavailable (absolute) or maxUnavailablePercentage — exactly one of the two | Absolute values bounded by a 100-node quota; percentage values update in parallel "up to 100 nodes at once" | The scale-up phase launches the larger of 2×AZs or maxUnavailable new nodes before draining — a 5-AZ group with maxUnavailable=1 launches up to 10 new nodes; maxUnavailable=20 launches 20 |
| Karpenter | disruption.budgets (nodes: "10%" default) | Most restrictive active budget wins | terminationGracePeriod explicitly "bypasses any blocked PDBs or the karpenter.sh/do-not-disrupt annotation" once the deadline hits — the escape hatch that exists precisely because PDB-blocked drains otherwise never finish |
| Cluster Autoscaler | No user-facing pacing; per-node eviction retries | Retries evictions for up to 2 minutes, then saves the node and moves on | During those 2 minutes, "other CA activity is stopped" (FAQ) — a blocked PDB stalls not just the drain but the autoscaler's other work |
Notice the pattern across all three: none of them negotiates with the PDB — they either wait, cap parallelism, or (Karpenter) break the budget with an explicit deadline. The missing primitive is the thing KEP-4563 is building.
The other two clouds expose the same knobs with different names, and it pays to know all three dialects. AKS node pool upgrades are paced by maxSurge/maxUnavailable plus a node drain timeout and a soak time between nodes — Microsoft's own guidance tunes "surge, PDB, node drain timeout, and node soak time" together to raise the odds of low-disruption upgrades, and suggests rounding surge to a multiple of three nodes to keep zone balance. GKE's surge upgrades follow the same shape: maxSurge provisioned first, maxUnavailable bounding the parallel drain. And if you drive EKS from code, eksctl exposes the same updateConfig (maxUnavailable or maxUnavailablePercentage, one or the other) per nodegroup in your manifest — the GitOps path for a setting the console hides.
Architectural Blueprint: The EvictionRequest API (KEP-4563)
Alpha in Kubernetes 1.37 behind the EvictionRequestAPI feature gate (kube-apiserver and kube-controller-manager), the new EvictionRequest resource in lifecycle.k8s.io/v1alpha1 turns eviction from a synchronous request-response into a declarative, observable, multi-requester coordination object. The API reference's own words: an EvictionRequest "should ideally result in a graceful eviction of a .spec.target," and the evictionrequest-controller "observes intents of all EvictionRequests and transforms them into Evictions."
The spec is deliberately small. Three fields, all required, two immutable:
apiVersion: lifecycle.k8s.io/v1alpha1
kind: EvictionRequest
metadata:
name: drain-node-42-pg-7f9c
namespace: payments # must match the target pod's namespace
spec:
intent: Eviction # "Eviction" | "Withdrawn" — enum, mutable
requester: drain.example.com/node-roller # domain-prefixed, immutable, required
target: # EvictionRequestTarget — immutable, required
pod: # EvictionRequestPodReference
name: pg-7f9c-xyz
uid: 8e0f6a44-…-… # lowercase UUID from target .metadata.uidThe semantics that matter for drain automation, straight from the reference:
- Many-to-many request/target. Multiple requesters (your drainer, the descheduler, a human) can request the same pod; the controller folds them into one Eviction. For a pod target the relationship is many-to-one: all requests collapse into one actual eviction.
- Withdrawal cancels. Flip
intent: Withdrawn(or delete the object). "If all requesters withdraw their eviction intent for a common target, the eviction will be canceled" — inactive responders never run, active ones are expected to cancel, completed ones take no action. This is the first built-in primitive for "abort the wave." - Responder model. Each target can have responders assigned; responders observe Evictions and implement the eviction logic, updating progress in the Eviction's status. The KEP's Alpha2 milestone proposes the responder registration mechanism.
- Status is conditions, not a return code.
TargetEvictedandFailedconditions (managed by the evictionrequest-controller) replace the 429-or-200 binary.Failedcovers canceled requests and "no responder managed to evict the target" — and conditions can be reset when a new Eviction intent is submitted. That is a machine-readable drain audit trail, which the 429 storm never gave you.
And the flow that replaces the serial kubectl loop:
sequenceDiagram
participant R as Drainer (node-roller)
participant API as kube-apiserver
participant C as evictionrequest-controller
participant W as Workload owner (responder)
R->>API: POST EvictionRequest (intent=Eviction)
API->>C: reconcile observed intents
C->>API: create/patch Eviction for target pod
W->>C: observes Eviction, decides when it is safe
W-->>C: updates progress in Eviction status
C->>R: condition TargetEvicted / Failed
R->>API: intent Withdrawn (abort wave) — eviction canceledWhy this fixes the three standing problems:
- The 429 storm becomes a condition you can watch. Instead of N parallel drainers polling a starved budget and eating API server bandwidth, each files its intent once and watches the object. Blocked is a state, not a retry loop — killing the retry storm from #106286 at the source.
- "Evict when you can" finally exists. The workload (via responders) participates in when the eviction executes. The KEP's stated mission is "cooperative termination of a pod, usually in order to run the pod on another node" — the exact semantics a fleet drainer needs and PDBs never had.
- Multi-requester fleet drains become composable. Karpenter filing an EvictionRequest and your upgrade tool filing another for the same pod is a feature of the many-to-many model, not a race. Who wins, and what happens on withdrawal, is now API-defined instead of "whichever client's retry loop hits the admission check first."
Do not rush the timeline, though. It is alpha: the feature gate is off by default, and the KEP's rollback section is blunt — disabling the gate kills pending eviction requests, the affected pods keep running, and any component built on the API stops working. Treat it as design-target material for your internal tooling roadmap, not a switch to flip on the prod control plane this quarter.
Production Code: A Fleet Drain Orchestrator That Survives PDBs
Until EvictionRequest ships stable, fleet drains are yours to orchestrate. Here is a complete, runnable pattern for a 1,000-node fleet: PDB-aware wave planning, per-node drain with real deadlines, Prometheus-sourced budget back-pressure, and a stuck-drain detector. Bash and YAML only — the parts you can actually run with tools you already have.
1. Fleet wave planner (Bash + jq + kubectl)
#!/usr/bin/env bash
# fleet-drain-waves.sh — plan bounded drain waves for a node list.
# Usage: fleet-drain-waves.sh <nodefile> <max_parallel> <drain_timeout>
# nodefile: one node name per line (from kubectl get nodes -o name)
# max_parallel: simultaneous in-flight drains (respect your PDB headroom;
# for EKS managed groups this must stay ≤ updateConfig.maxUnavailable)
set -euo pipefail
NODEFILE="${1:?node file required}"
MAX_PARALLEL="${2:?max parallel required}"
DRAIN_TIMEOUT="${3:?drain timeout per node, e.g. 15m}"
nodes=()
total=0
while IFS= read -r n; do
nodes+=("$n")
total=$((total+1))
done < "$NODEFILE"
# Waves of max_parallel nodes. Every node in a wave drains concurrently;
# the wave is not done until ALL its nodes finish or time out.
wave=0
for ((i=0; i<total; i+=MAX_PARALLEL)); do
wave=$((wave+1))
nslice=0
slice=()
for ((j=i; j<total && j<i+MAX_PARALLEL; j++)); do
slice+=("${nodes[j]}")
nslice=$((nslice+1))
done
printf 'wave %03d: %d nodes :: %s\n' "$wave" "$nslice" "${slice[*]}"
done
printf 'total: %d nodes, %d per wave, %s per-node timeout\n' \
"$total" "$MAX_PARALLEL" "$DRAIN_TIMEOUT"Two design notes. First, max_parallel is a PDB property, not a comfort setting: if your most constrained shared PDB allows 1 disruption and 10 nodes drain in parallel, at most one of those nodes holds that workload's pod at a time — the other nine wait on the same budget, and you have not actually parallelized anything. Second, the per-node timeout must be finite (--timeout 15m, never the default 0s infinite wait) so a wedged eviction surfaces as a failed node in the wave report instead of an invisible hang.
2. PDB headroom gate (PromQL before every wave)
# Disruption headroom across the cluster: how many evictions the PDBs
# will admit right now. Gate the wave on this, not on hope.
sum(kube_poddisruptionbudget_status_pod_disruptions_allowed)
# Per-PDB drill-down for the workloads you are about to disrupt:
kube_poddisruptionbudget_status_pod_disruptions_allowed
/ on (namespace, poddisruptionbudget) group_left
kube_poddisruptionbudget_status_expected_pods
# The trap detector: PDBs that will block EVERY eviction.
# disruptions_allowed = 0 while expected_pods > 0 → dead budget.
count(kube_poddisruptionbudget_status_pod_disruptions_allowed == 0
and on (namespace, poddisruptionbudget)
kube_poddisruptionbudget_status_expected_pods > 0)
# Capacity canary: cordoned-but-occupied nodes are silent capacity loss.
sum(kube_node_spec_unschedulable and on(node) kube_node_status_condition{condition="Ready"})The metric names are the kube-state-metrics PDB set — all STABLE tier except annotations/labels/deletion-timestamp, which you almost certainly do not need. The zero-headroom count is the pre-flight that catches the "PDB with no expected pods blocks everything" trap before it eats your window.
3. Per-node drain with deadline, force escalation, and report
#!/usr/bin/env bash
# drain-one.sh — drain a single node with a hard deadline and audit trail.
# Usage: drain-one.sh <node> <timeout> <force_after_s>
# Design: two-stage drain. Stage 1 respects PDBs with a finite timeout.
# Stage 2 (--force) is a HUMAN decision, surfaced as data, never automatic.
set -uo pipefail
NODE="${1:?node required}"
TIMEOUT="${2:-15m}"
FORCE_AFTER="${3:-0}" # seconds to wait before REPORTING force candidates
log() { printf '%s [%s] %s\n' "$(date -u +%FT%TZ)" "$NODE" "$*"; }
# Cordon first — even if the drain fails, the node must not take new work.
kubectl cordon "$NODE" || { log "cordon failed"; exit 2; }
log "stage 1: PDB-respecting drain, timeout ${TIMEOUT}"
if kubectl drain "$NODE" \
--ignore-daemonsets \
--delete-emptydir-data \
--timeout="$TIMEOUT"; then
log "stage 1 clean"
exit 0
fi
# Stage 1 failed: PDB blocked or a pod never terminated. Collect the evidence.
log "stage 1 failed — collecting blocked pods"
kubectl get pods --all-namespaces --field-selector "spec.nodeName=${NODE}" \
-o json | jq -r '.items[]
| select(.metadata.ownerReferences[]?.kind != "DaemonSet")
| "\(.metadata.namespace)/\(.metadata.name)\t\(.metadata.ownerReferences[0].kind)\t\(.status.phase)"' \
| tee "${NODE}.blocked"
log "blocked pods saved to ${NODE}.blocked — review before any --force"
log "force candidates (pods with no controlling owner):"
grep -v $'\t(Deployment|StatefulSet|DaemonSet|Job)\t' "${NODE}.blocked" || true
exit 1The two-stage discipline matters more than the script. --disable-eviction exists in the CLI surface and it does exactly what it sounds like — bypasses PDB checks entirely. Keep it in the human runbook, out of the automation. The escalation path that is safe to automate is Karpenter's, and it's worth copying its contract: a node-level terminationGracePeriod that waits out PDBs and do-not-disrupt until the deadline, then force-deletes — with the pre-alignment behavior (pods whose own grace period would straddle the deadline get deleted early at T = node timeout − pod terminationGracePeriodSeconds) so nothing is SIGKILLed mid-flight. That behavior is documented directly in the Karpenter API docs and the NodePool API source, and it is the best-available template for "bounded wait, then enforced drain" until EvictionRequest grows up.
4. The NodePool config that makes Karpenter's pacing explicit
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general-purpose
spec:
disruption:
# Budgets are how Karpenter paces ITS drains. Default if unset: one
# budget of 10%. Most restrictive active budget wins. Schedule+duration
# pairs give you business-hours protection.
budgets:
- nodes: "10%" # steady-state ceiling on terminating nodes
- nodes: "0"
schedule: "0 9 * * mon-fri"
duration: 8h # freeze consolidation during work hours
consolidateAfter: 5m
consolidationPolicy: WhenEmptyOrUnderutilized
template:
spec:
# TerminationGracePeriod is the bounded-wait contract: drains respect
# PDBs and do-not-disrupt until this deadline, then force-delete.
# WARNING (from the API docs): it bypasses blocked PDBs once reached.
terminationGracePeriod: 30m
nodeClassRef:
group: ec2.karpenter.sh
kind: EC2NodeClass
name: default5. Node "health" for drain eligibility: conditions, not vibes
#!/usr/bin/env bash
# drain-candidates.sh — nodes worth draining, in priority order.
# Priority: unhealthy > drift > underutilized. Never drain by age alone.
set -euo pipefail
# Unhealthy first — the fleet's real fires. (Karpenter 1.x routes these
# through the same disruption budgets as consolidation via reason Unhealthy.)
kubectl get nodes -o json | jq -r '.items[]
| select(
(.status.conditions[]? | select(.type=="Ready" and .status=="False")) or
(.status.conditions[]? | select(.type=="MemoryPressure" and .status=="True")) or
(.status.conditions[]? | select(.type=="DiskPressure" and .status=="True"))
)
| .metadata.name' > /tmp/candidates-unhealthy.txt
# Then drift: nodes not running the current golden AMI / image version.
# providerID is the stable per-cloud node identifier (e.g. aws:///i-abc123);
# match it against your golden launch-template data source of choice.
CURRENT_IMAGE="ami-your-golden-2026.09"
kubectl get nodes -o json | jq -r '.items[]
| .metadata.name + " " + (.spec.providerID // "unknown")' \
> /tmp/candidates-drift-full.txt || true
# join against your golden-AMI mapping (e.g. aws ec2 describe-instances)
# and keep only the nodes still on an old image:
# comm -23 <(sort /tmp/candidates-drift-full.txt) <(sort golden.txt) > /tmp/candidates-drift.txt
wc -l /tmp/candidates-unhealthy.txt /tmp/candidates-drift-full.txtDay-2 Operational Warnings
Metrics to watch. The drain fleet's health is a closed loop: PDB headroom (gating), cordoned nodes (stalled capacity), evictions admitted vs. blocked (throughput). The PromQL above covers the first two. For the third, watch the PDB gauges themselves — and when the EvictionRequest API eventually reaches your clusters, the KEP gives you a purpose-built counter: imperative_eviction_responder_controller_failed_evictions_total, incremented once after the first three failed attempts, with the pod requeued on exponential backoff capped around 15 minutes. But remember the docs' warning that 429s are ambiguous between budget-block and API throttling, so alert on the PDB headroom gauge directly rather than inferring from 429 rates alone.
Runbook notes. Three failure signatures and what they mean:
- Drain stuck at scale, one node, forever — check
kubectl get pdb -AforDISRUPTIONS: 0/Nagainst the stuck pod's workloads; check whether the replacement pods are failing probes (the docs' named cause: replacements never Ready → eviction 429s "until you intervene"). - Many nodes cordoned, capacity shrinking — that is your wave report going red; abort the wave (uncordon un-drained nodes), fix the budget, re-plan. Never let a wedged wave silently eat capacity.
- 429s from everything, not just evictions — API rate limiting, not budgets; the eviction endpoint shares that status code with throttling. Check API Priority and Fairness, not PDBs.
Blast radius. A drain wave gone wrong is a capacity incident, not an availability incident — until it is. Cordoned nodes hold pods that keep serving; the moment you force past a PDB you have traded availability for the upgrade window, and the PDB was the only thing in the room with the arithmetic to push back. The multi-PDB deadlock (#75957) and the preemption-vs-PDB violation (#91492) both mean: never treat "PDB blocks me" as an application bug to override — it is the system working exactly as designed, and the override path belongs to a human with a deadline and a name.
References & Further Reading
- API-initiated Eviction — Kubernetes documentation — the Eviction API, 429 semantics, and the stuck-eviction troubleshooting section quoted throughout.
- Specifying Disruptions for your Application — Kubernetes documentation — PDB status fields, the disruptions-allowed formula, percentage rounding, and unhealthyPodEvictionPolicy.
- EvictionRequest API reference (lifecycle.k8s.io/v1alpha1) — the complete v1alpha1 spec: intent, requester, target, and the TargetEvicted/Failed conditions.
- KEP-4563: EvictionRequest API — kubernetes/enhancements — sig-node; motivation, standing PDB issues, responder model, rollback semantics.
- kubectl drain reference — the full flag surface including --disable-eviction, --timeout, and --skip-wait-for-delete-seconds.
- Issue #114877: maxSurge for node draining — the open ask for capacity-first drain pacing.
- Issue #75957: cannot drain node with pods matched by multiple PDBs — the multi-PDB deadlock.
- Issue #91492: preemption can violate PDBs — the standing defect class KEP-4563 aims to design out.
- Issue #106286: confusing use of TooManyRequests for eviction — why your runbook cannot tell budget-block from throttling.
- Karpenter — GitHub (kubernetes-sigs) — disruption budgets and the terminationGracePeriod bounded-wait contract (NodePool v1 API).
- Cluster Autoscaler FAQ — per-node eviction retry window (2 minutes) and PDB respect in scale-down.
- Amazon EKS managed node group update behavior — updateConfig.maxUnavailable/maxUnavailablePercentage, the 100-node quota, and the scale-up formula.
- Graceful node shutdown — Kubernetes documentation — kubelet shutdownGracePeriod and the two-phase shutdown ordering.
- kube-state-metrics PDB metrics — the STABLE-tier gauges used in the PromQL back-pressure gate.