Kubernetes Node Drains at Scale: PDB Math, 429 Storms, and the EvictionRequest API That Fixes Both

Sources

Nobody writes a blog post about the day the drain worked. The stories that reach your incident channel are the other kind: an upgrade wave that was supposed to finish in two hours and is still wedged at node 412 of 640 eight hours later, with a postgres pod that refuses to move and a PDB that says DisruptionsAllowed: 0 with total sincerity. This guide is about why that happens in 1,000+ node clusters specifically — the arithmetic, the client behaviors, and the controller mechanics that are invisible at 40 nodes and load-bearing at 1,000 — and what the new EvictionRequest API (KEP-4563, alpha in Kubernetes 1.37) actually changes.

First, the part everyone gets wrong: a node drain is not a kubectl feature — it is a loop over the Eviction API, and the Eviction API is a request-response endpoint over a PodDisruptionBudget admission check. kubectl drain is a serial client walking that endpoint one pod at a time, honoring PDBs, honoring graceful termination, and never — by design — telling the API server "I need this whole node empty by 02:00." The official eviction docs are candid about the consequence: when a pod's replacement never becomes Ready, the eviction keeps returning 429 Too Many Requests "until you intervene." At 40 nodes, intervention is an SRE with a terminal. At 1,000 nodes, intervention is a job title.

The Structural Shift: From "Delete the Pods" to "Coordinated, Budgeted Disruption"

The original Kubernetes answer to "empty this node" was kubectl delete pod --all-namespaces --field-selector spec.nodeName=... — destructive, PDB-blind, and fast. Everything since has been a long retreat from that position, adding layers of coordination between your intent ("this node goes away at 02:00") and the pod's fate. Count the layers that a modern drain passes through:

your intent: "node-7 dies tonight"
        │
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ drain orchestrator (kubectl / Karpenter / CA / EKS updater)      │
│   decides WHICH node, WHEN, and how to pace parallel drains     │
└─────────────────────────────────────────────────────────────────┘
        │  POST /api/v1/namespaces/ns/pods/NAME/eviction
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ kube-apiserver: Eviction API (policy/v1 subresource)            │
│   • admission: find PDBs matching the pod                       │
│   • if budget would be violated → 429 TooManyRequests           │
│   • success → pod deletionTimestamp set, graceful shutdown clock│
└─────────────────────────────────────────────────────────────────┘
        │
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ PodDisruptionBudget controller (kube-controller-manager)         │
│   recomputes DisruptionsAllowed every sync from live pod health │
└─────────────────────────────────────────────────────────────────┘
        │  terminationGracePeriodSeconds counts down
        ▼
┌─────────────────────────────────────────────────────────────────┐
│ kubelet on the old node: preStop hooks → SIGTERM → SIGKILL      │
│   (Graceful Node Shutdown adds two-phase ordering on reboots)   │
└─────────────────────────────────────────────────────────────────┘
        │  endpoints/EndpointSlice controller removes the IP
        ▼
   traffic drains; scheduler places the replacement elsewhere

Every arrow in that diagram is a place where a 1,000-node drain dies. Not the arrows themselves — the coordination gaps between them, which is exactly the framing the KEP-4563 authors use: today's tooling drives API-initiated eviction "in an application agnostic way," and PDBs — the application's only standard defense — were never designed to express "evict this pod when you can, not now or never."

Layer 1: PDB math is status math, not spec math

The budget that stalls your drain is computed from live cluster state, not from the YAML you wrote. Per the PDB documentation, the status fields are:

currentHealthy    3   ← number of pods reporting Ready (live)
desiredHealthy    2   ← spec-derived: expectedPods - allowed disruptions
expectedPods      3   ← total pods matched by .spec.selector
disruptionsAllowed 1 ← currentHealthy - desiredHealthy
observedGeneration 1

The formula from the same docs: disruptionsAllowed = currentHealthy - desiredHealthy. That is the entire gate. Note what is not in it: your deadline, your wave plan, your upgrade window. The PDB does not know a drain is happening. It only knows how many pods are Ready right now — which is why a drain that would be trivially fine at steady state becomes wedged when the replacement pods from the previous wave haven't passed their probes yet.

Two edge cases from the same page bite harder at scale. Percentage values: minAvailable: "50%" with 7 pods rounds up to 4 — so small workloads with percentage PDBs lose more budget than you'd guess. And a maxUnavailable: 0% with zero expected pods produces DisruptionsAllowed: 0 — the classic "PDB covers zero pods and blocks everything" trap. The docs' own example of this: a single-instance app where the PDB admits no disruption at all; their stated solution is to use minAvailable: "90%" instead, or drop the PDB entirely.

Layer 2: kubectl drain is a serial client, and serial does not scale

kubectl drain walks the node's pods one eviction at a time. It is a one-node tool with honest defaults, and the generated reference shows its full flag surface:

--disable-eviction=false     Force drain to use delete, bypassing PDB checks
--delete-emptydir-data=false  Continue even if pods use emptyDir
--force=false                Continue even if pods have no controller
--grace-period=-1            Use the pod's own terminationGracePeriodSeconds
--ignore-daemonsets=false    Ignore DaemonSet-managed pods
--pod-selector=''            Filter pods by label selector
--skip-wait-for-delete-timeout=0  Skip waiting for pods already deleting >N s
--timeout=0s                 Give up after this long; zero means infinite

Read --timeout 0s — zero means infinite again, because that is the default posture of the tool most teams reach for during an incident: it will sit under a blocked eviction forever, and the docs warn explicitly that evictions can be stuck "if the last evicted Pod had a long termination grace period" or if new pods never become Ready. None of these flags let you say "pace this across the fleet" or "abort this wave at 04:00." Those are orchestrator concerns, and the orchestrators each answer them differently — EKS caps parallel node drains at 10, Karpenter budgets them by percentage, Cluster Autoscaler retries for 2 minutes and then moves on.

The Numbers That Break at 1,000 Nodes

At 40 nodes, drain failure is a rounding error. At 1,000, every per-pod constant multiplies into a wall-clock problem. The three that matter:

ConstantValueWhere it comes fromAt 1,000 nodes
Serial eviction per node1 pod at a timekubectl drain client loopPods per node × (PDB wait + grace period) stacks up per node, serially
PDB budget refresh cadencedisruption controller synckube-controller-managerA blocked PDB stays blocked until replacements become Ready — no back-pressure to the drain
Eviction API response on violation429 TooManyRequestsAPI-initiated eviction docsRetry storms across parallel drainers multiply API server load

Now multiply it out with a worked example. This is the arithmetic of a stalled upgrade wave — run the same numbers for your own fleet:

# Assumptions (verify against your own cluster):
#   1000 nodes, 10 pods/node worth evicting, PDB allows 1 disruption per workload
#   kubectl drain --timeout 15m per node (no infinite wait), serial within node
#   replacement pods need 120s to become Ready (probe delays)
#
# Steady state, everything healthy:
#   per pod: eviction accepted → grace period 30s → replacement Ready at +120s
#   per node: 10 pods, but PDB allows 1 at a time → ~10 × 150s = 25 min/node
#   PDB refill: next disruption allowed once replacement becomes Ready
#
# The wedge: any pod whose replacement never becomes Ready
#   → eviction returns 429 forever (docs: "until you intervene")
#   → kubectl drain --timeout expires → node stays cordoned, still occupied
#   → cordoned nodes never rejoin rotation → capacity shrinks silently
#
# Second-order: parallel drainers all polling the same starved PDBs
#   → 429 storm on the API server (also indistinguishable from rate limiting)
#   → your automation sees "429" and cannot tell budget-block from throttle

The last line is the operational kicker the KEP authors called out directly: issue #106286 — "Confusing use of TooManyRequests error for eviction" — is one of the standing PDB issues the KEP lists. Your runbook cannot distinguish "the application is unhealthy" from "you're hitting API priority and fairness limits" because the API uses the same status code for both.

The multi-PDB deadlock

The other classic wedge: one pod selected by two PDBs. Per issue #75957 ("Cannot drain node with pod with more than one Pod Disruption Budget"), the eviction admission path checks each matching PDB and fails if any single budget would be violated — so overlapping selectors (say, a shared team=payments budget plus a per-app budget) can produce a pod that can never be evicted while any one of its budgets is tight, even when the other has budget to spare. The KEP's standing-issues list treats this as a defect class to design out, not a misconfiguration.

Where orchestrators already pace drains — and where they cap you

Nobody sane drains 1,000 nodes with kubectl. The real fleets use orchestrators, and their pacing is the actual "drain policy" of your cluster, whether you've read it or not:

OrchestratorDrain/pacing controlDefault / capGotcha
EKS managed node groupsupdateConfig.maxUnavailable (absolute) or maxUnavailablePercentage — exactly one of the twoAbsolute values bounded by a 100-node quota; percentage values update in parallel "up to 100 nodes at once"The scale-up phase launches the larger of 2×AZs or maxUnavailable new nodes before draining — a 5-AZ group with maxUnavailable=1 launches up to 10 new nodes; maxUnavailable=20 launches 20
Karpenterdisruption.budgets (nodes: "10%" default)Most restrictive active budget winsterminationGracePeriod explicitly "bypasses any blocked PDBs or the karpenter.sh/do-not-disrupt annotation" once the deadline hits — the escape hatch that exists precisely because PDB-blocked drains otherwise never finish
Cluster AutoscalerNo user-facing pacing; per-node eviction retriesRetries evictions for up to 2 minutes, then saves the node and moves onDuring those 2 minutes, "other CA activity is stopped" (FAQ) — a blocked PDB stalls not just the drain but the autoscaler's other work

Notice the pattern across all three: none of them negotiates with the PDB — they either wait, cap parallelism, or (Karpenter) break the budget with an explicit deadline. The missing primitive is the thing KEP-4563 is building.

The other two clouds expose the same knobs with different names, and it pays to know all three dialects. AKS node pool upgrades are paced by maxSurge/maxUnavailable plus a node drain timeout and a soak time between nodes — Microsoft's own guidance tunes "surge, PDB, node drain timeout, and node soak time" together to raise the odds of low-disruption upgrades, and suggests rounding surge to a multiple of three nodes to keep zone balance. GKE's surge upgrades follow the same shape: maxSurge provisioned first, maxUnavailable bounding the parallel drain. And if you drive EKS from code, eksctl exposes the same updateConfig (maxUnavailable or maxUnavailablePercentage, one or the other) per nodegroup in your manifest — the GitOps path for a setting the console hides.

Architectural Blueprint: The EvictionRequest API (KEP-4563)

Alpha in Kubernetes 1.37 behind the EvictionRequestAPI feature gate (kube-apiserver and kube-controller-manager), the new EvictionRequest resource in lifecycle.k8s.io/v1alpha1 turns eviction from a synchronous request-response into a declarative, observable, multi-requester coordination object. The API reference's own words: an EvictionRequest "should ideally result in a graceful eviction of a .spec.target," and the evictionrequest-controller "observes intents of all EvictionRequests and transforms them into Evictions."

The spec is deliberately small. Three fields, all required, two immutable:

apiVersion: lifecycle.k8s.io/v1alpha1
kind: EvictionRequest
metadata:
  name: drain-node-42-pg-7f9c
  namespace: payments      # must match the target pod's namespace
spec:
  intent: Eviction         # "Eviction" | "Withdrawn" — enum, mutable
  requester: drain.example.com/node-roller   # domain-prefixed, immutable, required
  target:                  # EvictionRequestTarget — immutable, required
    pod:                   # EvictionRequestPodReference
      name: pg-7f9c-xyz
      uid: 8e0f6a44-…-…    # lowercase UUID from target .metadata.uid

The semantics that matter for drain automation, straight from the reference:

And the flow that replaces the serial kubectl loop:

sequenceDiagram
    participant R as Drainer (node-roller)
    participant API as kube-apiserver
    participant C as evictionrequest-controller
    participant W as Workload owner (responder)
    R->>API: POST EvictionRequest (intent=Eviction)
    API->>C: reconcile observed intents
    C->>API: create/patch Eviction for target pod
    W->>C: observes Eviction, decides when it is safe
    W-->>C: updates progress in Eviction status
    C->>R: condition TargetEvicted / Failed
    R->>API: intent Withdrawn (abort wave) — eviction canceled

Why this fixes the three standing problems:

Do not rush the timeline, though. It is alpha: the feature gate is off by default, and the KEP's rollback section is blunt — disabling the gate kills pending eviction requests, the affected pods keep running, and any component built on the API stops working. Treat it as design-target material for your internal tooling roadmap, not a switch to flip on the prod control plane this quarter.

Production Code: A Fleet Drain Orchestrator That Survives PDBs

Until EvictionRequest ships stable, fleet drains are yours to orchestrate. Here is a complete, runnable pattern for a 1,000-node fleet: PDB-aware wave planning, per-node drain with real deadlines, Prometheus-sourced budget back-pressure, and a stuck-drain detector. Bash and YAML only — the parts you can actually run with tools you already have.

1. Fleet wave planner (Bash + jq + kubectl)

#!/usr/bin/env bash
# fleet-drain-waves.sh — plan bounded drain waves for a node list.
# Usage: fleet-drain-waves.sh <nodefile> <max_parallel> <drain_timeout>
#   nodefile: one node name per line (from kubectl get nodes -o name)
#   max_parallel: simultaneous in-flight drains (respect your PDB headroom;
#                 for EKS managed groups this must stay ≤ updateConfig.maxUnavailable)
set -euo pipefail

NODEFILE="${1:?node file required}"
MAX_PARALLEL="${2:?max parallel required}"
DRAIN_TIMEOUT="${3:?drain timeout per node, e.g. 15m}"

nodes=()
total=0
while IFS= read -r n; do
  nodes+=("$n")
  total=$((total+1))
done < "$NODEFILE"

# Waves of max_parallel nodes. Every node in a wave drains concurrently;
# the wave is not done until ALL its nodes finish or time out.
wave=0
for ((i=0; i<total; i+=MAX_PARALLEL)); do
  wave=$((wave+1))
  nslice=0
  slice=()
  for ((j=i; j<total && j<i+MAX_PARALLEL; j++)); do
    slice+=("${nodes[j]}")
    nslice=$((nslice+1))
  done
  printf 'wave %03d: %d nodes :: %s\n' "$wave" "$nslice" "${slice[*]}"
done
printf 'total: %d nodes, %d per wave, %s per-node timeout\n' \
  "$total" "$MAX_PARALLEL" "$DRAIN_TIMEOUT"

Two design notes. First, max_parallel is a PDB property, not a comfort setting: if your most constrained shared PDB allows 1 disruption and 10 nodes drain in parallel, at most one of those nodes holds that workload's pod at a time — the other nine wait on the same budget, and you have not actually parallelized anything. Second, the per-node timeout must be finite (--timeout 15m, never the default 0s infinite wait) so a wedged eviction surfaces as a failed node in the wave report instead of an invisible hang.

2. PDB headroom gate (PromQL before every wave)

# Disruption headroom across the cluster: how many evictions the PDBs
# will admit right now. Gate the wave on this, not on hope.
sum(kube_poddisruptionbudget_status_pod_disruptions_allowed)

# Per-PDB drill-down for the workloads you are about to disrupt:
kube_poddisruptionbudget_status_pod_disruptions_allowed
  / on (namespace, poddisruptionbudget) group_left
kube_poddisruptionbudget_status_expected_pods

# The trap detector: PDBs that will block EVERY eviction.
# disruptions_allowed = 0 while expected_pods > 0 → dead budget.
count(kube_poddisruptionbudget_status_pod_disruptions_allowed == 0
  and on (namespace, poddisruptionbudget)
  kube_poddisruptionbudget_status_expected_pods > 0)

# Capacity canary: cordoned-but-occupied nodes are silent capacity loss.
sum(kube_node_spec_unschedulable and on(node) kube_node_status_condition{condition="Ready"})

The metric names are the kube-state-metrics PDB set — all STABLE tier except annotations/labels/deletion-timestamp, which you almost certainly do not need. The zero-headroom count is the pre-flight that catches the "PDB with no expected pods blocks everything" trap before it eats your window.

3. Per-node drain with deadline, force escalation, and report

#!/usr/bin/env bash
# drain-one.sh — drain a single node with a hard deadline and audit trail.
# Usage: drain-one.sh <node> <timeout> <force_after_s>
# Design: two-stage drain. Stage 1 respects PDBs with a finite timeout.
# Stage 2 (--force) is a HUMAN decision, surfaced as data, never automatic.
set -uo pipefail

NODE="${1:?node required}"
TIMEOUT="${2:-15m}"
FORCE_AFTER="${3:-0}"   # seconds to wait before REPORTING force candidates

log() { printf '%s [%s] %s\n' "$(date -u +%FT%TZ)" "$NODE" "$*"; }

# Cordon first — even if the drain fails, the node must not take new work.
kubectl cordon "$NODE" || { log "cordon failed"; exit 2; }

log "stage 1: PDB-respecting drain, timeout ${TIMEOUT}"
if kubectl drain "$NODE" \
     --ignore-daemonsets \
     --delete-emptydir-data \
     --timeout="$TIMEOUT"; then
  log "stage 1 clean"
  exit 0
fi

# Stage 1 failed: PDB blocked or a pod never terminated. Collect the evidence.
log "stage 1 failed — collecting blocked pods"
kubectl get pods --all-namespaces --field-selector "spec.nodeName=${NODE}" \
  -o json | jq -r '.items[]
    | select(.metadata.ownerReferences[]?.kind != "DaemonSet")
    | "\(.metadata.namespace)/\(.metadata.name)\t\(.metadata.ownerReferences[0].kind)\t\(.status.phase)"' \
  | tee "${NODE}.blocked"

log "blocked pods saved to ${NODE}.blocked — review before any --force"
log "force candidates (pods with no controlling owner):"
grep -v $'\t(Deployment|StatefulSet|DaemonSet|Job)\t' "${NODE}.blocked" || true
exit 1

The two-stage discipline matters more than the script. --disable-eviction exists in the CLI surface and it does exactly what it sounds like — bypasses PDB checks entirely. Keep it in the human runbook, out of the automation. The escalation path that is safe to automate is Karpenter's, and it's worth copying its contract: a node-level terminationGracePeriod that waits out PDBs and do-not-disrupt until the deadline, then force-deletes — with the pre-alignment behavior (pods whose own grace period would straddle the deadline get deleted early at T = node timeout − pod terminationGracePeriodSeconds) so nothing is SIGKILLed mid-flight. That behavior is documented directly in the Karpenter API docs and the NodePool API source, and it is the best-available template for "bounded wait, then enforced drain" until EvictionRequest grows up.

4. The NodePool config that makes Karpenter's pacing explicit

apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: general-purpose
spec:
  disruption:
    # Budgets are how Karpenter paces ITS drains. Default if unset: one
    # budget of 10%. Most restrictive active budget wins. Schedule+duration
    # pairs give you business-hours protection.
    budgets:
      - nodes: "10%"           # steady-state ceiling on terminating nodes
      - nodes: "0"
        schedule: "0 9 * * mon-fri"
        duration: 8h            # freeze consolidation during work hours
    consolidateAfter: 5m
    consolidationPolicy: WhenEmptyOrUnderutilized
  template:
    spec:
      # TerminationGracePeriod is the bounded-wait contract: drains respect
      # PDBs and do-not-disrupt until this deadline, then force-delete.
      # WARNING (from the API docs): it bypasses blocked PDBs once reached.
      terminationGracePeriod: 30m
      nodeClassRef:
        group: ec2.karpenter.sh
        kind: EC2NodeClass
        name: default

5. Node "health" for drain eligibility: conditions, not vibes

#!/usr/bin/env bash
# drain-candidates.sh — nodes worth draining, in priority order.
# Priority: unhealthy > drift > underutilized. Never drain by age alone.
set -euo pipefail

# Unhealthy first — the fleet's real fires. (Karpenter 1.x routes these
# through the same disruption budgets as consolidation via reason Unhealthy.)
kubectl get nodes -o json | jq -r '.items[]
  | select(
      (.status.conditions[]? | select(.type=="Ready" and .status=="False")) or
      (.status.conditions[]? | select(.type=="MemoryPressure" and .status=="True")) or
      (.status.conditions[]? | select(.type=="DiskPressure" and .status=="True"))
    )
  | .metadata.name' > /tmp/candidates-unhealthy.txt

# Then drift: nodes not running the current golden AMI / image version.
# providerID is the stable per-cloud node identifier (e.g. aws:///i-abc123);
# match it against your golden launch-template data source of choice.
CURRENT_IMAGE="ami-your-golden-2026.09"
kubectl get nodes -o json | jq -r '.items[]
  | .metadata.name + " " + (.spec.providerID // "unknown")' \
  > /tmp/candidates-drift-full.txt || true
# join against your golden-AMI mapping (e.g. aws ec2 describe-instances)
# and keep only the nodes still on an old image:
# comm -23 <(sort /tmp/candidates-drift-full.txt) <(sort golden.txt) > /tmp/candidates-drift.txt

wc -l /tmp/candidates-unhealthy.txt /tmp/candidates-drift-full.txt

Day-2 Operational Warnings

Metrics to watch. The drain fleet's health is a closed loop: PDB headroom (gating), cordoned nodes (stalled capacity), evictions admitted vs. blocked (throughput). The PromQL above covers the first two. For the third, watch the PDB gauges themselves — and when the EvictionRequest API eventually reaches your clusters, the KEP gives you a purpose-built counter: imperative_eviction_responder_controller_failed_evictions_total, incremented once after the first three failed attempts, with the pod requeued on exponential backoff capped around 15 minutes. But remember the docs' warning that 429s are ambiguous between budget-block and API throttling, so alert on the PDB headroom gauge directly rather than inferring from 429 rates alone.

Runbook notes. Three failure signatures and what they mean:

Blast radius. A drain wave gone wrong is a capacity incident, not an availability incident — until it is. Cordoned nodes hold pods that keep serving; the moment you force past a PDB you have traded availability for the upgrade window, and the PDB was the only thing in the room with the arithmetic to push back. The multi-PDB deadlock (#75957) and the preemption-vs-PDB violation (#91492) both mean: never treat "PDB blocks me" as an application bug to override — it is the system working exactly as designed, and the override path belongs to a human with a deadline and a name.

References & Further Reading