Kueue v0.20.0 Upgrade Field Guide: The Release That Retires v1beta1
Sources
- Kueue v0.20.0 release notes (published 2026-09-30)
- Kueue source tree, tag v0.20.0 — pkg/features/kube_features.go (feature-gate registry and defaults)
- Kueue source tree, tag v0.20.0 — apis/config/v1beta2/defaults.go (WaitForPodsReady defaulting)
- Kueue docs: Installation (k8s 1.34+ recommended, Helm OCI install)
- Kueue docs: Admission Fair Sharing (accounting anchor at quota reservation, v0.20)
- Kueue docs: Set up Wait for Pods Ready
- Kueue docs: Prometheus metrics reference
- Kueue docs: Set up Dynamic Resource Allocation (k8s 1.34+)
- Kueue docs: ClusterQueue concepts (flavors, borrowing, preemption policies)
- Kueue: migrate-to-v1beta2.sh migration script (hack/)
- Kueue issue #13320 — FairSharingPrioritizeNonBorrowing leaf-CQ starvation (open)
- Kueue PR #13863 — RecomputeAssignmentUponPreemptionTargetsOverlap (merged 2026-08-11)
- Kueue issue #14543 — FairSharing infinite preemption loop (open)
- Kueue PR #13279 — LWSImmutableGroupSize quota-bypass fix
- Kueue PR #15833 — SparkApplication cores/memoryOverhead quota fix
- Kubernetes docs: API deprecation policy
Kueue v0.20.0 landed on 30 September 2026, and it is not a routine minor. The Kueue maintainers use the .0 of each minor to graduate feature gates and retire API surface, and this one does both at once: kueue.x-k8s.io/v1beta1 is no longer served — the API directory still contains the types, but only as conversion stubs marked +kubebuilder:unservedversion, so any client, GitOps manifest, or CI pipeline still pinned to v1beta1 breaks on the spot. Around that headline sit twelve feature gates that silently flipped to Beta-on, four quota-accounting bypasses that got fixed (meaning your quotas were quietly wrong until now), and a fair-sharing accounting change that alters admission order for any ClusterQueue using UsageBasedAdmissionFairSharing.
This is a field guide, not a release-notes paraphrase. Every gate name, default value, and API field below was verified against the v0.20.0 source tree — the feature-gate registry in pkg/features/kube_features.go, the config defaulting in apis/config/v1beta2/defaults.go, and the scheduler's admission path in pkg/scheduler/scheduler.go. If you run Kueue in production — and in 2026, if you run AI training or batch on Kubernetes, you either run Kueue or something shaped like it — the sequence below is what gets you from v0.19 to v0.20 without an incident review.
TL;DR
- The API break is real and immediate: v0.20.0 stops serving
kueue.x-k8s.io/v1beta1entirely. Storage moved to v1beta2 back in v0.16; serving is now v1beta2-only. Run the migration script before the upgrade, not after, and audit every pipeline that kubectl-applies Kueue YAML. - Twelve gates flipped to Beta, enabled by default — including
WorkloadPriorityClassDefaulting(every unlabeled workload now gets thedefaultpriority class),AdmissionFairSharingAnchorAtQuotaReservation(fair-sharing accounting starts at quota reservation, not admission), andLWSImmutableGroupSize(LeaderWorkerSet group size becomes immutable while Kueue-managed). - Four quota bypasses closed — LWS group-size scaling, Spark
cores/memoryOverheadunder-reservation, negative-request quota credit, and Job parallelism-over-completions early release. If your quotas "felt tight" before, they were lying: these bugs let workloads run beyond quota. - WaitForPodsReady is on by default since v0.19 (30-minute timeout, 30-minute recovery). If you upgraded 0.18 → 0.19 and skipped that note, v0.20 is your forced reckoning: the escape hatch gate
DisableWaitForPodsReadyis deprecated with removal planned in 0.21. - Upgrade order that works: pre-flight audit → migration script → canary ClusterQueue → controller roll → watch
kueue_cluster_queue_statusand AFS metrics → unbake v1beta1 from CI. Detailed runbook below.
If you want the concepts rather than the migration — how priority, fair sharing, cohorts, and DRA quota fit together for LLM fleets — read our GPU scheduling for LLM workloads guide first; this piece assumes you know what a ClusterQueue is and picks up at the version you are now running.
The Structural Shift: This Is a Versioned-Kubernetes-API Retirement, Not a Config Bump
Most Kueue upgrades are boring. You bump the Helm chart, the controller image rolls, CRDs apply server-side, done. v0.20.0 is different for one structural reason: Kueue's CRDs are now single-serving-version. The Kubernetes API deprecation policy governs how served versions are retired, and Kueue's v1beta1 has reached the end of that runway. Concretely, from the v0.20.0 source tree: apis/kueue/v1beta1/ still exists, but every type file carries +kubebuilder:unservedversion — the generated CRD no longer lists v1beta1 in spec.versions[].served, so the API server itself rejects requests for it. This is not a deprecation warning; it is a hard 404 on the version string.
Storage vs serving matters here. Since v0.16.0, new objects are stored as v1beta2, and the release notes of that cycle told everyone to run the migration script. The script (hack/migrate-to-v1beta2.sh) performs a no-op update on every existing Kueue object, forcing each through the conversion webhook so etcd ends up holding v1beta2. Objects never touched by that script sit in etcd as v1beta1 storage with a conversion webhook keeping them readable — until the serving version goes away and the conversion path breaks with it. v0.20.0 is the release where that debt comes due: an unmigrated object becomes unreadable to any v1beta2-only client, and the migration script can no longer list objects via v1beta1 by default (its --kubectl-get-api-version flag exists for exactly this stranded state, defaulting to v1beta1 for clusters "just after upgrade to Kueue 0.16+" — read the script's help before assuming).
The second structural shift is quieter and, in practice, more disruptive: the default-on gate graduation batch. Kueue's pattern is to flip alpha gates to Beta-on in .0 releases. Verified from pkg/features/kube_features.go in the v0.20.0 tree, the gates that became Default: true, PreRelease: Beta in v0.20 include:
AdmissionFairSharingAnchorAtQuotaReservation # AFS accounting anchor moves to quota reservation
DeploymentJobUIDLabel # Workload of Deployment-managed Pod carries Deployment UID
FlavorFungibilityPreserveScanProgress # retry workloads across flavors, not just the first one
HighMaxParallelismWithinReconcile # 8 -> 32 concurrent API updates per reconcile
LWSImmutableGroupSize # leaderWorkerTemplate.size immutable while Kueue-managed
MultiKueueKubeConfigPathValidation # kubeconfig path refs restricted to /etc/multikueue/kubeconfigs/
MultiKueueOrchestratedPreemption # coordinated preemption across worker clusters
MultiKueueRemoteSpecSync # in-place Ray Serve config updates forwarded to workers
MultiKueueReuseClientConnectionConfigForWorkers # QPS/Burst reuse for worker-cluster clients
PodGroupSchedulingShapeOrdering # deterministic flavor assignment for Pod groups
PodIntegrationCountSucceededPodsAsReady # succeeded pods count as ready for PodsReady
RecomputeAssignmentUponPreemptionTargetsOverlap # hero-workload starvation fix, in-cycle recompute
RejectUpdatesToCQWithInvalidOnFlavors # reject CQ updates referencing invalid flavor names
SkipAncestorCheckForDeletedWorkloads # GC not blocked by "workload owner not found"
SchedulingEquivalenceHashingIgnorePodSetName # PodSet names ignored in equivalence hashing
TASCacheTopologyTree # cached topology trees between cycles
TASGroupedPodSetSlicing # LWS leader grouped + workers sliced
TASNodeFeasibilityForAllLevels # per-node feasibility at every topology level
TASPartialSlices # PodSet counts not a multiple of slice size
UnadmittedWorkloadsObservability # granular pending reasons in QuotaReserved condition
WorkloadPriorityClassDefaulting # default WorkloadPriorityClass auto-assigned
WorkloadValidateResourcesAreNonNegative # negative resource requests rejected at validation
WorkloadValidationForPodSetMetadata # invalid PodSet template metadata rejected at admission
EnforceProvisioningPodTemplateContents # divergent PodTemplate specs replaced with Kueue-derivedTwenty-three gates. Most are pure hygiene (a rejection that previously arrived as a stuck-pending workload now arrives as a validation error — better, and usually invisible). But four of them change steady-state behavior in ways an on-call rotation will notice:
WorkloadPriorityClassDefaulting— every workload that never specified a priority class now gets one. If your queueing strategy isStrictFIFO, the default class keeps relative order; if you rely on unlabeled workloads sorting below explicitly-classed ones, that assumption just died.AdmissionFairSharingAnchorAtQuotaReservation— for ClusterQueues withadmissionScope.admissionMode: UsageBasedAdmissionFairSharing, fair-sharing usage now accrues from quota reservation, not admission. A workload sitting in AdmissionChecks for an hour now depresses its LocalQueue's fair-share standing for that hour. Teams with slow AdmissionChecks (external provisioning, cluster-autoscaler ProvisioningRequest flows) will see admission order shift.LWSImmutableGroupSize—spec.leaderWorkerTemplate.sizeon a Kueue-managed LeaderWorkerSet is now immutable. This is the quota-bypass fix (see below); the operational consequence is that autoscaling an LWS's group size now means delete-and-recreate, not in-place edit.spec.replicasstays mutable — scale the number of groups, not the size of each.TASNodeFeasibilityForAllLevels— topologies whose lowest level is notkubernetes.io/hostnamenow check node feasibility per node. Workloads that used to fit a domain's aggregate capacity but not any single node are no longer admitted. If your TAS queues show newPendingworkloads post-upgrade with inadmissible reasons about node selectors or taints, this gate is why.
Also note the deprecation: DisableWaitForPodsReady (introduced in v0.19.0 as the escape hatch for WaitForPodsReady becoming default-on) is marked "planned to be removed in 0.21" in the source. If you run with this gate set to true, start planning the config-side change now — and note v0.20 adds WaitForPodsReadyUnscheduledTimeout (alpha, off) and WorkloadLevelWaitForPodsReady (alpha, off), the first knobs for per-workload timeout overrides via the kueue.x-k8s.io/wait-for-pods-ready annotation (JSON: {"timeoutSeconds": ..., "recoveryTimeoutSeconds": ...}, bounded by waitForPodsReady.maxTimeoutOnWorkload, default cap 2h from the source defaults).
Architectural Blueprint: Where Each Landmine Sits in the Admission Path
To sequence the upgrade correctly you need to see where each breaking change executes. The diagram below traces a workload from submission through the v0.20.0 admission path, with the four behavior changes marked at their point of effect.
+---------------------------+
Job/CR submit | Mutating & validating |
-----------------> | webhooks |
+------------+--------------+
| [1] v1beta1 manifests rejected here
| [2] default WorkloadPriorityClass stamped here
v
+----------------------------+
| Workload object created |
| (scheduling-gated) |
+------------+---------------+
| queued via LocalQueue -> ClusterQueue
v
+--------------------------------------------+
| SCHEDULER CYCLE (snapshot + one pass) |
| |
| 1. equivalence-hash cluster (Scheduling- |
| EquivalenceHashing: same-shape work- |
| loads share evaluation) |
| 2. flavor assignment (flavorassigner) |
| 3. [3] AFS: LocalQueue usage + entry |
| penalty now anchored at QUOTA |
| RESERVATION, not admission |
| 4. [4] preemption candidate recompute on |
| target overlap (RecomputeAssignment- |
| UponPreemptionTargetsOverlap) |
| 5. admission (or inadmissible + reason) |
+---------------------+----------------------+
|
+----------------------+----------------------+
| |
QuotaReserved=True inadmissible
(usage now charged [3]) (UnadmittedWorkloadsObservability:
| granular reason in QuotaReserved
v condition + metrics)
+---------------------+ AdmissionChecks (ProvisioningRequest,
| Pods ungated; | MultiKueue dispatch, DRA feasibility)
| kube-scheduler | |
| places pods | v
+----------+----------+ Admitted=True
| |
v v
WaitForPodsReady (default since v0.19: 30m timeout,
30m recovery; Disable gate dies in 0.21)
|
v
TAS topologies: per-node feasibility checks
(TASNodeFeasibilityForAllLevels) + partial slicesRead the diagram with the upgrade sequence in mind: changes [1] and [2] fire at webhook time and break things loudly; [3] and [4] fire inside the scheduler cycle and shift outcomes silently. Loud breaks you fix during the upgrade window; silent shifts you detect with the metrics in the Day-2 section.
The Upgrade Runbook
Assumptions: Helm-managed Kueue (chart 0.20.0, appVersion v0.20.0), Kubernetes 1.34+ per the installation docs (v0.20's own test matrix runs baseline E2E on 1.34–1.37), and a cluster where Kueue already runs v0.16+ so storage is v1beta2 or convertible. If you are on something older, stop: upgrade to v0.19.x first, run the migration, stabilize, then proceed.
Step 0 — Pre-flight audit (before touching anything)
# 1. Which Kueue objects still exist, and does anything still speak v1beta1?
kubectl get workloads,localqueues,clusterqueues,resourceflavors,cohorts \
--all-namespaces -o custom-columns=KIND:.kind,NAME:.metadata.name --no-headers \
| sort | uniq -c | sort -rn
# 2. Where does v1beta1 still appear in your GitOps repos and CI?
# (grep every repo that applies Kueue manifests)
grep -rn "kueue.x-k8s.io/v1beta1" argo/ flux/ .github/ gitlab-ci/ 2>/dev/null
# 3. Is anything using an external scheduler or pre-0.16 client library?
grep -rn "kueue.x-k8s.io" --include="*.go" --include="*.py" services/ 2>/dev/null \
| grep -v v1beta2
# 4. Snapshot current queue state for post-upgrade comparison
kubectl get clusterqueues -o json > /tmp/cq-pre-upgrade.json
kubectl get localqueues -A -o json > /tmp/lq-pre-upgrade.json
kubectl get workloads -A -o json > /tmp/wl-pre-upgrade.json
# 5. Record the running controller's effective config (Helm values or ConfigMap)
kubectl -n kueue-system get configmap kueue-manager-config -o yaml \
> /tmp/kueue-config-pre.yamlStep 0.2 is the one teams skip and regret. Any pipeline that kubectl applys a manifest with apiVersion: kueue.x-k8s.io/v1beta1 will fail after the upgrade — not with a deprecation warning, with no matches for kind / the server could not find the requested resource, because the version is simply no longer in the CRD's served list. Fix the manifests before the upgrade, not during the incident.
Step 1 — Run the v1beta2 migration script against the old controller
# Download the migration script (from the main branch; behavior matches the release tree)
wget -O /tmp/migrate-to-v1beta2.sh \
https://raw.githubusercontent.com/kubernetes-sigs/kueue/main/hack/migrate-to-v1beta2.sh
chmod +x /tmp/migrate-to-v1beta2.sh
# Dry-run FIRST — the script supports client and server dry-run
/tmp/migrate-to-v1beta2.sh --dry-run=server
# Review the output: it lists every object it will touch, per kind and namespace.
# Only when the dry-run is clean:
/tmp/migrate-to-v1beta2.sh
# Verify: every Workload should now report storageVersion v1beta2
kubectl get workloads -A -o jsonpath='{range .items[*]}{.metadata.name}{"\n"}{end}' \
| head -20The script's own defaults: it lists objects via v1beta1 by default (flag --kubectl-get-api-version, default v1beta1, intended for clusters "just after upgrade to Kueue 0.16+"), patches them with a no-op update to force conversion, and supports --namespace-regex to work in batches on large clusters. On a cluster still running v0.19, v1beta1 is still served, so the script can list and patch everything. Do this before the controller upgrade — after the upgrade, the default listing version may find nothing to migrate.
Step 2 — Pre-stage the v0.20.0 chart and diff the CRDs
# Pull the release chart (OCI registry path per the official docs)
helm pull oci://registry.k8s.io/kueue/charts/kueue --version 0.20.0 --untar -d /tmp/kueue-chart
# CRDs first — this is where the v1beta1 serving removal is visible
ls /tmp/kueue-chart/kueue/crds/
grep -A3 "served:" /tmp/kueue-chart/kueue/crds/kueue.x-k8s.io_clusterqueues.yaml | head -20
# Expect: only v1beta2 with served: true
# Render your full values against the new chart and diff against the live cluster
helm template kueue /tmp/kueue-chart/kueue -n kueue-system \
-f your-kueue-values.yaml > /tmp/kueue-020-rendered.yaml
kubectl diff -f /tmp/kueue-020-rendered.yamlStep 3 — Decide the feature-gate posture explicitly
The v0.20 chart exposes feature gates as a list under controllerManager.featureGates, rendered into the manager's --feature-gates flag. Because so many gates flipped on, the safest production posture for one release cycle is to pin the four behavior-changing ones off, upgrade, then enable them one per day while watching the Day-2 metrics:
controllerManager:
featureGates:
# Behavior-changing v0.20 Beta graduations — enable one per day post-roll:
- name: WorkloadPriorityClassDefaulting
enabled: false
- name: AdmissionFairSharingAnchorAtQuotaReservation
enabled: false
- name: LWSImmutableGroupSize
enabled: false
- name: TASNodeFeasibilityForAllLevels
enabled: false
managerConfig:
controllerManagerConfigYaml: |
apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
health:
healthProbeBindAddress: :8081
leaderElection:
leaderElect: true
waitForPodsReady:
timeout: 30m
recoveryTimeout: 30m
# NOTE: since v0.19 an ABSENT block means the defaults (enabled, 30m).
# To DISABLE, set the DisableWaitForPodsReady feature gate —
# and plan for its 0.21 removal.Careful with the gate dependency the release notes flag: AdmissionFairSharingAnchorAtQuotaReservation depends on AdmissionFairSharing and is enabled by default; disabling AdmissionFairSharing alone is rejected at startup unless you also disable AdmissionFairSharingAnchorAtQuotaReservation in the same configuration. The docs are explicit on this pairing, and the source encodes the dependency in its gate-dependency map. If you run AFS at all, decide on the pair explicitly in values.yaml rather than inheriting defaults.
Step 4 — Canary the upgrade on one ClusterQueue
Kueue upgrades are controller-wide — you cannot canary the binary itself — but you can canary the blast radius. Before the roll, create a sacrificial queue pair and route only a test team to it:
apiVersion: kueue.x-k8s.io/v1beta2
kind: ResourceFlavor
metadata:
name: canary-spot
spec:
nodeTaints: []
nodeSelector:
node.kubernetes.io/instance-type: spot-small
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: cq-canary
spec:
namespaceSelector:
matchLabels:
kueue-canary: "true"
queueingStrategy: BestEffortFIFO
admissionScope:
admissionMode: UsageBasedAdmissionFairSharing
resourceGroups:
- coveredResources: ["cpu", "memory"]
flavors:
- name: canary-spot
resources:
- name: cpu
nominalQuota: 100
- name: memory
nominalQuota: 400Gi
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: LocalQueue
metadata:
name: lq-canary
namespace: team-canary
spec:
clusterQueue: cq-canary
fairSharing:
weight: "1"Apply this after the controller roll and watch one full workload lifecycle through it: QuotaReserved → AdmissionChecks → Admitted → PodsReady. The QuotaReserved=False condition with the new UnadmittedWorkloadsObservability gate gives you a machine-readable reason — WaitingForQuota, ExceedsMaxQuota, WaitingForPodsReady, Misconfigured, or Suspended — instead of a bare Pending count. That single change makes canary triage dramatically faster than in v0.19.
Step 5 — Roll the controller and watch for the three failure signatures
# The roll (per the official installation docs)
helm upgrade kueue oci://registry.k8s.io/kueue/charts/kueue \
--version 0.20.0 \
--namespace kueue-system \
-f your-kueue-values.yaml \
--wait --timeout 300s
# SIGNATURE 1: startup rejection — a gate dependency you didn't know about
kubectl -n kueue-system logs deploy/kueue-controller-manager | grep -i "feature.gate\|reject"
# e.g. "AdmissionFairSharingAnchorAtQuotaReservation requires AdmissionFairSharing"
# Fix: correct the featureGates list; do NOT blind-restart.
# SIGNATURE 2: webhook panics — the v0.20 webhook is stricter
kubectl -n kueue-system logs deploy/kueue-controller-manager | grep -i "panic\|webhook"
# New rejections at admission: invalid PodSet metadata (WorkloadValidationForPodSetMetadata),
# negative resource requests (WorkloadValidateResourcesAreNonNegative),
# TAS slice-size validation (PodSetSliceSize >= 1).
# SIGNATURE 3: stuck-pending workloads that used to run
kubectl get workloads -A -o json \
| jq -r '.items[] | select(.status.conditions[]? | select(.type=="QuotaReserved" and .status=="False")) | .metadata.name'
# Check .status.conditions[].reason — if it names node selectors/taints you
# never configured, TASNodeFeasibilityForAllLevels is reclassifying capacity.The Four Quota Bypasses (Why Your Quotas Were Wrong Before)
Four fixes in v0.20.0 close accounting holes that let workloads consume more than their ClusterQueue quota covered. These are not theoretical: each shipped as a fix with an issue number, and each means your observed utilization pre-upgrade was understated. Expect ClusterQueue usage to look "higher" after the upgrade not because anything changed, but because the ledger finally matches reality.
| Bypass fixed | Mechanism before v0.20 | Fix in v0.20.0 |
|---|---|---|
| LeaderWorkerSet group-size scaling (#13279) | Raising spec.leaderWorkerTemplate.size on an admitted, Kueue-managed LWS ran more pods per group than reserved quota covered | leaderWorkerTemplate.size immutable while Kueue-managed (LWSImmutableGroupSize, Beta-on). Recreate at new size instead; spec.replicas (group count) stays mutable |
| Spark under-reservation (#15833) | Workloads reserved less CPU/memory than Spark actually requested: cores, memoryOverhead, and Spark's default memory overhead were ignored | Each Pod now reserves its cores as CPU plus at least 384Mi additional memory. Action: raise ClusterQueue quotas for Spark queues on upgrade — every Spark driver/executor got cheaper-to-count, not more expensive to run |
| Negative-request quota credit (#12838) | Negative container resource requests created artificial ClusterQueue quota credit, bypassing configured limits | Negative values floored to zero in accounting, rejected at validation by default (WorkloadValidateResourcesAreNonNegative, Beta-on) |
| Job parallelism early release (#16111) | A Job with parallelism > completions released the quota of still-running Pods once some Pods succeeded, admitting beyond quota | Quota held until Pods terminate, not until siblings succeed |
The Spark fix deserves emphasis because it is the one that requires a quota change on upgrade day. Kueue now reserves Spark's cores as CPU and adds the default memory overhead — so a ClusterQueue sized for the old accounting will start rejecting Spark workloads that previously fit. Memory values must use Spark's Java format (512m, 2g); Kubernetes-style 512Mi is rejected. If you run Spark on Kubernetes under Kueue, budget the quota bump before the roll.
The AFS Accounting Anchor: Who Gets Deprioritized Now
For ClusterQueues running UsageBasedAdmissionFairSharing — the admission-time fairness layer, distinct from the older cohort preemption fair sharing — v0.20 moves the accounting anchor. Verified in pkg/scheduler/scheduler.go (only Usage-based ClusterQueues settle penalties; the anchor change applies at quota reservation) and the AFS docs:
- Before: a workload's fair-share usage contribution (and entry penalty) started when it was admitted. Time spent waiting in AdmissionChecks was free.
- After: it starts when quota is reserved. A workload can hold quota while its ProvisioningRequest or MultiKueue dispatch is pending — and that hold now counts against its LocalQueue's fair-share standing.
Who wins and loses: LocalQueues whose workloads pass AdmissionChecks quickly are unaffected. LocalQueues whose workloads sit in checks for a long time — external provisioning waits, cross-cluster dispatch — now accumulate usage for that waiting time and get deprioritized relative to leaner queues. If a team complains that "our jobs stopped getting in after the upgrade," this anchor move is the first place to look: check kubectl get localqueue <name> -o jsonpath={.status.fairSharing} — the consumedResources and weightedShare fields are the live evidence.
Two open issues temper the fair-sharing story, and honesty requires naming them: #14543 (FairSharing infinite preemption loop, open) is the known failure mode the FairSharingReevaluatePreemptionCandidates alpha fix risks aggravating, and #13320 documents a leaf-CQ non-borrowing starvation case. The v0.20 PrioritizePreemptorWorkloads alpha gate addresses preemption thrash, disabled by default. Fair sharing in Kueue is powerful but still has sharp edges: if you run it, read both issues before the upgrade, and keep the preemption-strategy config (LessThanOrEqualToFinalShare / LessThanInitialShare) pinned in review.
Day-2: Metrics, Alerts, and Blast Radius
The metrics that catch a sick queue
All names below verified against pkg/metrics/metrics.go in the v0.20.0 source (68 registered metrics); the rendered reference is in the Kueue metrics docs. The v0.20 additions worth adding to dashboards:
# Fair-share usage per LocalQueue — watch this shift after the anchor move
kueue_local_queue_admission_fair_sharing_usage
# Unadmitted workloads, by granular reason (UnadmittedWorkloadsObservability, now on)
kueue_unadmitted_workloads
# Pending workloads per ClusterQueue — the classic health signal
kueue_pending_workloads
# NEW v0.20: time from admission to PodsReady recovery, per the recovery flow
kueue_workload_recovery_wait_time_seconds
# NEW v0.20: distinct pending scheduling-equivalence hashes per ClusterQueue
kueue_pending_scheduling_hashes
# NEW v0.20: overlapping preemption target recomputations in-cycle
kueue_preemption_target_recomputations_total
# NEW v0.20: MultiKueue worker-cluster connectivity status
kueue_multikueue_cluster_status
# NEW v0.20: workload total execution time across admissions
kueue_execution_time_seconds
# NEW v0.20: evictions on worker clusters by cluster + reason
kueue_multikueue_workloads_evicted_totalAlerting deltas to configure on upgrade day:
| Alert | Expression (starting point) | Why v0.20 changed it |
|---|---|---|
| AFS usage anomaly | deriv(kueue_local_queue_admission_fair_sharing_usage[1h]) > 0 sustained on previously-quiet LocalQueues | The quota-reservation anchor starts charging queues whose workloads wait in AdmissionChecks |
| Unexpected pending spike | kueue_unadmitted_workloads > 0 on queues that ran clean pre-upgrade | TASNodeFeasibilityForAllLevels reclassifies domain-fit workloads as inadmissible; Spark accounting raises per-pod reservations |
| Preemption churn | rate(kueue_preempted_workloads_total[15m]) > 0 new baseline | RecomputeAssignmentUponPreemptionTargetsOverlap changes preemption timing; #14543's loop risk lives here |
| Recovery wait regression | histogram_quantile(0.95, kueue_workload_recovery_wait_time_seconds_bucket) rising | New metric — baseline it now, alert on drift later |
Runbook notes
- Rollback path: downgrade the chart to 0.19.7. But note the config asymmetry the release notes call out: if you configured
AdmissionFairSharingAnchorAtQuotaReservationexplicitly, remove it from the configuration before rolling back to a release that does not recognize it, or the older controller rejects its config at startup. The same applies to any explicitly-set v0.20 gate name. - Do not skip .0 releases in one jump. The upgrade notes for v0.20 explicitly require reviewing the
.0notes for each minor you cross (v0.18.0, v0.19.0). The WaitForPodsReady default-on, the DRA gate renames, and the AdmissionGatedBy graduation all live in those notes. - Blast radius of a bad upgrade: the controller manager is leader-elected and single-replica by default; a crash-looping manager leaves existing admitted workloads running (pods are ungated) but freezes all new admissions and evictions. Your batch fleet keeps draining; your queue keeps growing. This is why the canary queue in Step 4 exists — admission freezes are invisible until someone's job doesn't start.
- KueueViz got real RBAC in v0.20 (#13810): the dashboard backend now performs SubjectAccessReviews per view when authentication is enabled. Non-Helm installs need the backend ServiceAccount to be able to create
subjectaccessreviewsinauthorization.k8s.io— the Helm chart grants it; hand-rolled RBAC does not. If KueueViz panels go blank post-upgrade, that permission is the first thing to check. - Logger names changed (#15259): core controllers now log as
<kind>-reconciler(e.g.clusterqueue-reconciler) and subcomponents as<subcomponent>-<kind>-reconciler(e.g.multikueue-workload-reconciler). Any log-based alert or Loki pipeline matching on the old names silently matches nothing after the upgrade.
Should You Upgrade This Week?
- Yes, if you run LWS at scale (the group-size bypass was real overcommit), run Spark under Kueue (your ledger was lying), or need the per-workload WaitForPodsReady annotations coming behind the alpha gates. The migration script path is well-trodden and the failure signatures are loud.
- Wait one patch release, if you are heavy on MultiKueue + RayService (the zero-downtime upgrade path
RayServiceValidateUpgradeStrategyBeta-on plus theElasticJobsViaWorkloadSlicesannotation requirement is a behavior change at upgrade time), or if you cannot yet purge v1beta1 from CI — because the API break turns that from tech debt into a page. - Skip the auto-bump, always: Kueue's
.0releases are where the sharp edges ship by design. Pin 0.20.0 in week one if you upgrade early, watch the release feed for 0.20.1 (the project's cadence publishes patch releases from active branches — v0.19.7 and v0.18.11 shipped the same day as v0.20.0), and upgrade to the patch once it lands.
References & Further Reading
- Kueue v0.20.0 release notes — the primary source: API removal, gate graduations, and all four quota-bypass fixes with issue numbers.
- pkg/features/kube_features.go @ v0.20.0 — the feature-gate registry and default table this guide's gate claims were verified against.
- apis/config/v1beta2/defaults.go @ v0.20.0 — config defaulting: WaitForPodsReady, 30-minute timeout defaults, 2-hour maxTimeoutOnWorkload cap.
- hack/migrate-to-v1beta2.sh @ v0.20.0 — the migration script: dry-run modes, namespace regex batching, kubectl API-version flags.
- Kueue installation docs — Helm OCI install path and the Kubernetes 1.34+ recommendation.
- Kueue docs: Admission Fair Sharing — the accounting-anchor change, entry penalty, and the gate dependency.
- Kueue docs: Wait for Pods Ready setup — the feature's config surface, default-on since v0.19.
- Kueue docs: Prometheus metrics reference — rendered tables of all exported metrics.
- Kueue docs: Set up Dynamic Resource Allocation — DRA integration requirements (Kubernetes 1.34+) and the KueueDRA* gate set.
- Kueue docs: ClusterQueue concepts — flavors, nominalQuota, borrowing and lending limits, preemption policies.
- Kubernetes API deprecation policy — the policy that governed the v1beta1 retirement timeline.
- Issue #14543: FairSharing infinite preemption loop and #13320: leaf-CQ non-borrowing starvation — the open fair-sharing sharp edges, read before enabling.
- Red Hat OpenShift AI: distributed workloads with Kueue — the commercial distribution path: the Red Hat build of Kueue operator, CodeFlare SDK, and KubeRay integration.
- GKE: Ray + Kueue + Dynamic Workload Scheduler tutorial — Google's managed integration of Kueue with DWS capacity queues.
- Kueue project documentation and the kubernetes-sigs/kueue repository.