Kueue v0.20.0 Upgrade Field Guide: The Release That Retires v1beta1

Sources

Kueue v0.20.0 landed on 30 September 2026, and it is not a routine minor. The Kueue maintainers use the .0 of each minor to graduate feature gates and retire API surface, and this one does both at once: kueue.x-k8s.io/v1beta1 is no longer served — the API directory still contains the types, but only as conversion stubs marked +kubebuilder:unservedversion, so any client, GitOps manifest, or CI pipeline still pinned to v1beta1 breaks on the spot. Around that headline sit twelve feature gates that silently flipped to Beta-on, four quota-accounting bypasses that got fixed (meaning your quotas were quietly wrong until now), and a fair-sharing accounting change that alters admission order for any ClusterQueue using UsageBasedAdmissionFairSharing.

This is a field guide, not a release-notes paraphrase. Every gate name, default value, and API field below was verified against the v0.20.0 source tree — the feature-gate registry in pkg/features/kube_features.go, the config defaulting in apis/config/v1beta2/defaults.go, and the scheduler's admission path in pkg/scheduler/scheduler.go. If you run Kueue in production — and in 2026, if you run AI training or batch on Kubernetes, you either run Kueue or something shaped like it — the sequence below is what gets you from v0.19 to v0.20 without an incident review.

TL;DR

If you want the concepts rather than the migration — how priority, fair sharing, cohorts, and DRA quota fit together for LLM fleets — read our GPU scheduling for LLM workloads guide first; this piece assumes you know what a ClusterQueue is and picks up at the version you are now running.

The Structural Shift: This Is a Versioned-Kubernetes-API Retirement, Not a Config Bump

Most Kueue upgrades are boring. You bump the Helm chart, the controller image rolls, CRDs apply server-side, done. v0.20.0 is different for one structural reason: Kueue's CRDs are now single-serving-version. The Kubernetes API deprecation policy governs how served versions are retired, and Kueue's v1beta1 has reached the end of that runway. Concretely, from the v0.20.0 source tree: apis/kueue/v1beta1/ still exists, but every type file carries +kubebuilder:unservedversion — the generated CRD no longer lists v1beta1 in spec.versions[].served, so the API server itself rejects requests for it. This is not a deprecation warning; it is a hard 404 on the version string.

Storage vs serving matters here. Since v0.16.0, new objects are stored as v1beta2, and the release notes of that cycle told everyone to run the migration script. The script (hack/migrate-to-v1beta2.sh) performs a no-op update on every existing Kueue object, forcing each through the conversion webhook so etcd ends up holding v1beta2. Objects never touched by that script sit in etcd as v1beta1 storage with a conversion webhook keeping them readable — until the serving version goes away and the conversion path breaks with it. v0.20.0 is the release where that debt comes due: an unmigrated object becomes unreadable to any v1beta2-only client, and the migration script can no longer list objects via v1beta1 by default (its --kubectl-get-api-version flag exists for exactly this stranded state, defaulting to v1beta1 for clusters "just after upgrade to Kueue 0.16+" — read the script's help before assuming).

The second structural shift is quieter and, in practice, more disruptive: the default-on gate graduation batch. Kueue's pattern is to flip alpha gates to Beta-on in .0 releases. Verified from pkg/features/kube_features.go in the v0.20.0 tree, the gates that became Default: true, PreRelease: Beta in v0.20 include:

AdmissionFairSharingAnchorAtQuotaReservation   # AFS accounting anchor moves to quota reservation
DeploymentJobUIDLabel                            # Workload of Deployment-managed Pod carries Deployment UID
FlavorFungibilityPreserveScanProgress           # retry workloads across flavors, not just the first one
HighMaxParallelismWithinReconcile               # 8 -> 32 concurrent API updates per reconcile
LWSImmutableGroupSize                          # leaderWorkerTemplate.size immutable while Kueue-managed
MultiKueueKubeConfigPathValidation              # kubeconfig path refs restricted to /etc/multikueue/kubeconfigs/
MultiKueueOrchestratedPreemption                # coordinated preemption across worker clusters
MultiKueueRemoteSpecSync                       # in-place Ray Serve config updates forwarded to workers
MultiKueueReuseClientConnectionConfigForWorkers # QPS/Burst reuse for worker-cluster clients
PodGroupSchedulingShapeOrdering                 # deterministic flavor assignment for Pod groups
PodIntegrationCountSucceededPodsAsReady         # succeeded pods count as ready for PodsReady
RecomputeAssignmentUponPreemptionTargetsOverlap # hero-workload starvation fix, in-cycle recompute
RejectUpdatesToCQWithInvalidOnFlavors           # reject CQ updates referencing invalid flavor names
SkipAncestorCheckForDeletedWorkloads            # GC not blocked by "workload owner not found"
SchedulingEquivalenceHashingIgnorePodSetName    # PodSet names ignored in equivalence hashing
TASCacheTopologyTree                            # cached topology trees between cycles
TASGroupedPodSetSlicing                         # LWS leader grouped + workers sliced
TASNodeFeasibilityForAllLevels                  # per-node feasibility at every topology level
TASPartialSlices                                # PodSet counts not a multiple of slice size
UnadmittedWorkloadsObservability                # granular pending reasons in QuotaReserved condition
WorkloadPriorityClassDefaulting                 # default WorkloadPriorityClass auto-assigned
WorkloadValidateResourcesAreNonNegative        # negative resource requests rejected at validation
WorkloadValidationForPodSetMetadata             # invalid PodSet template metadata rejected at admission
EnforceProvisioningPodTemplateContents          # divergent PodTemplate specs replaced with Kueue-derived

Twenty-three gates. Most are pure hygiene (a rejection that previously arrived as a stuck-pending workload now arrives as a validation error — better, and usually invisible). But four of them change steady-state behavior in ways an on-call rotation will notice:

Also note the deprecation: DisableWaitForPodsReady (introduced in v0.19.0 as the escape hatch for WaitForPodsReady becoming default-on) is marked "planned to be removed in 0.21" in the source. If you run with this gate set to true, start planning the config-side change now — and note v0.20 adds WaitForPodsReadyUnscheduledTimeout (alpha, off) and WorkloadLevelWaitForPodsReady (alpha, off), the first knobs for per-workload timeout overrides via the kueue.x-k8s.io/wait-for-pods-ready annotation (JSON: {"timeoutSeconds": ..., "recoveryTimeoutSeconds": ...}, bounded by waitForPodsReady.maxTimeoutOnWorkload, default cap 2h from the source defaults).

Architectural Blueprint: Where Each Landmine Sits in the Admission Path

To sequence the upgrade correctly you need to see where each breaking change executes. The diagram below traces a workload from submission through the v0.20.0 admission path, with the four behavior changes marked at their point of effect.

                    +---------------------------+
   Job/CR submit     |  Mutating & validating     |
----------------->  |  webhooks                  |
                    +------------+--------------+
                                 |  [1] v1beta1 manifests rejected here
                                 |  [2] default WorkloadPriorityClass stamped here
                                 v
                    +----------------------------+
                    |  Workload object created   |
                    |  (scheduling-gated)        |
                    +------------+---------------+
                                 |  queued via LocalQueue -> ClusterQueue
                                 v
              +--------------------------------------------+
              |  SCHEDULER CYCLE (snapshot + one pass)     |
              |                                            |
              |  1. equivalence-hash cluster (Scheduling-   |
              |     EquivalenceHashing: same-shape work-   |
              |     loads share evaluation)                |
              |  2. flavor assignment (flavorassigner)     |
              |  3. [3] AFS: LocalQueue usage + entry       |
              |     penalty now anchored at QUOTA           |
              |     RESERVATION, not admission              |
              |  4. [4] preemption candidate recompute on   |
              |     target overlap (RecomputeAssignment-    |
              |     UponPreemptionTargetsOverlap)          |
              |  5. admission (or inadmissible + reason)    |
              +---------------------+----------------------+
                                    |
             +----------------------+----------------------+
             |                                             |
   QuotaReserved=True                              inadmissible
   (usage now charged [3])                  (UnadmittedWorkloadsObservability:
             |                               granular reason in QuotaReserved
             v                               condition + metrics)
  +---------------------+   AdmissionChecks (ProvisioningRequest,
  | Pods ungated;       |   MultiKueue dispatch, DRA feasibility)
  | kube-scheduler      |        |
  | places pods         |        v
  +----------+----------+   Admitted=True
             |                  |
             v                  v
      WaitForPodsReady (default since v0.19: 30m timeout,
        30m recovery; Disable gate dies in 0.21)
             |
             v
      TAS topologies: per-node feasibility checks
        (TASNodeFeasibilityForAllLevels) + partial slices

Read the diagram with the upgrade sequence in mind: changes [1] and [2] fire at webhook time and break things loudly; [3] and [4] fire inside the scheduler cycle and shift outcomes silently. Loud breaks you fix during the upgrade window; silent shifts you detect with the metrics in the Day-2 section.

The Upgrade Runbook

Assumptions: Helm-managed Kueue (chart 0.20.0, appVersion v0.20.0), Kubernetes 1.34+ per the installation docs (v0.20's own test matrix runs baseline E2E on 1.34–1.37), and a cluster where Kueue already runs v0.16+ so storage is v1beta2 or convertible. If you are on something older, stop: upgrade to v0.19.x first, run the migration, stabilize, then proceed.

Step 0 — Pre-flight audit (before touching anything)

# 1. Which Kueue objects still exist, and does anything still speak v1beta1?
kubectl get workloads,localqueues,clusterqueues,resourceflavors,cohorts \
  --all-namespaces -o custom-columns=KIND:.kind,NAME:.metadata.name --no-headers \
  | sort | uniq -c | sort -rn

# 2. Where does v1beta1 still appear in your GitOps repos and CI?
#    (grep every repo that applies Kueue manifests)
grep -rn "kueue.x-k8s.io/v1beta1" argo/ flux/ .github/ gitlab-ci/ 2>/dev/null

# 3. Is anything using an external scheduler or pre-0.16 client library?
grep -rn "kueue.x-k8s.io" --include="*.go" --include="*.py" services/ 2>/dev/null \
  | grep -v v1beta2

# 4. Snapshot current queue state for post-upgrade comparison
kubectl get clusterqueues -o json > /tmp/cq-pre-upgrade.json
kubectl get localqueues -A -o json > /tmp/lq-pre-upgrade.json
kubectl get workloads -A -o json > /tmp/wl-pre-upgrade.json

# 5. Record the running controller's effective config (Helm values or ConfigMap)
kubectl -n kueue-system get configmap kueue-manager-config -o yaml \
  > /tmp/kueue-config-pre.yaml

Step 0.2 is the one teams skip and regret. Any pipeline that kubectl applys a manifest with apiVersion: kueue.x-k8s.io/v1beta1 will fail after the upgrade — not with a deprecation warning, with no matches for kind / the server could not find the requested resource, because the version is simply no longer in the CRD's served list. Fix the manifests before the upgrade, not during the incident.

Step 1 — Run the v1beta2 migration script against the old controller

# Download the migration script (from the main branch; behavior matches the release tree)
wget -O /tmp/migrate-to-v1beta2.sh \
  https://raw.githubusercontent.com/kubernetes-sigs/kueue/main/hack/migrate-to-v1beta2.sh
chmod +x /tmp/migrate-to-v1beta2.sh

# Dry-run FIRST — the script supports client and server dry-run
/tmp/migrate-to-v1beta2.sh --dry-run=server

# Review the output: it lists every object it will touch, per kind and namespace.
# Only when the dry-run is clean:
/tmp/migrate-to-v1beta2.sh

# Verify: every Workload should now report storageVersion v1beta2
kubectl get workloads -A -o jsonpath='{range .items[*]}{.metadata.name}{"\n"}{end}' \
  | head -20

The script's own defaults: it lists objects via v1beta1 by default (flag --kubectl-get-api-version, default v1beta1, intended for clusters "just after upgrade to Kueue 0.16+"), patches them with a no-op update to force conversion, and supports --namespace-regex to work in batches on large clusters. On a cluster still running v0.19, v1beta1 is still served, so the script can list and patch everything. Do this before the controller upgrade — after the upgrade, the default listing version may find nothing to migrate.

Step 2 — Pre-stage the v0.20.0 chart and diff the CRDs

# Pull the release chart (OCI registry path per the official docs)
helm pull oci://registry.k8s.io/kueue/charts/kueue --version 0.20.0 --untar -d /tmp/kueue-chart

# CRDs first — this is where the v1beta1 serving removal is visible
ls /tmp/kueue-chart/kueue/crds/
grep -A3 "served:" /tmp/kueue-chart/kueue/crds/kueue.x-k8s.io_clusterqueues.yaml | head -20
# Expect: only v1beta2 with served: true

# Render your full values against the new chart and diff against the live cluster
helm template kueue /tmp/kueue-chart/kueue -n kueue-system \
  -f your-kueue-values.yaml > /tmp/kueue-020-rendered.yaml
kubectl diff -f /tmp/kueue-020-rendered.yaml

Step 3 — Decide the feature-gate posture explicitly

The v0.20 chart exposes feature gates as a list under controllerManager.featureGates, rendered into the manager's --feature-gates flag. Because so many gates flipped on, the safest production posture for one release cycle is to pin the four behavior-changing ones off, upgrade, then enable them one per day while watching the Day-2 metrics:

controllerManager:
  featureGates:
    # Behavior-changing v0.20 Beta graduations — enable one per day post-roll:
    - name: WorkloadPriorityClassDefaulting
      enabled: false
    - name: AdmissionFairSharingAnchorAtQuotaReservation
      enabled: false
    - name: LWSImmutableGroupSize
      enabled: false
    - name: TASNodeFeasibilityForAllLevels
      enabled: false
  managerConfig:
    controllerManagerConfigYaml: |
      apiVersion: config.kueue.x-k8s.io/v1beta2
      kind: Configuration
      health:
        healthProbeBindAddress: :8081
      leaderElection:
        leaderElect: true
      waitForPodsReady:
        timeout: 30m
        recoveryTimeout: 30m
        # NOTE: since v0.19 an ABSENT block means the defaults (enabled, 30m).
        # To DISABLE, set the DisableWaitForPodsReady feature gate —
        # and plan for its 0.21 removal.

Careful with the gate dependency the release notes flag: AdmissionFairSharingAnchorAtQuotaReservation depends on AdmissionFairSharing and is enabled by default; disabling AdmissionFairSharing alone is rejected at startup unless you also disable AdmissionFairSharingAnchorAtQuotaReservation in the same configuration. The docs are explicit on this pairing, and the source encodes the dependency in its gate-dependency map. If you run AFS at all, decide on the pair explicitly in values.yaml rather than inheriting defaults.

Step 4 — Canary the upgrade on one ClusterQueue

Kueue upgrades are controller-wide — you cannot canary the binary itself — but you can canary the blast radius. Before the roll, create a sacrificial queue pair and route only a test team to it:

apiVersion: kueue.x-k8s.io/v1beta2
kind: ResourceFlavor
metadata:
  name: canary-spot
spec:
  nodeTaints: []
  nodeSelector:
    node.kubernetes.io/instance-type: spot-small
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
  name: cq-canary
spec:
  namespaceSelector:
    matchLabels:
      kueue-canary: "true"
  queueingStrategy: BestEffortFIFO
  admissionScope:
    admissionMode: UsageBasedAdmissionFairSharing
  resourceGroups:
    - coveredResources: ["cpu", "memory"]
      flavors:
        - name: canary-spot
          resources:
            - name: cpu
              nominalQuota: 100
            - name: memory
              nominalQuota: 400Gi
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: LocalQueue
metadata:
  name: lq-canary
  namespace: team-canary
spec:
  clusterQueue: cq-canary
  fairSharing:
    weight: "1"

Apply this after the controller roll and watch one full workload lifecycle through it: QuotaReserved → AdmissionChecks → Admitted → PodsReady. The QuotaReserved=False condition with the new UnadmittedWorkloadsObservability gate gives you a machine-readable reason — WaitingForQuota, ExceedsMaxQuota, WaitingForPodsReady, Misconfigured, or Suspended — instead of a bare Pending count. That single change makes canary triage dramatically faster than in v0.19.

Step 5 — Roll the controller and watch for the three failure signatures

# The roll (per the official installation docs)
helm upgrade kueue oci://registry.k8s.io/kueue/charts/kueue \
  --version 0.20.0 \
  --namespace kueue-system \
  -f your-kueue-values.yaml \
  --wait --timeout 300s

# SIGNATURE 1: startup rejection — a gate dependency you didn't know about
kubectl -n kueue-system logs deploy/kueue-controller-manager | grep -i "feature.gate\|reject"
# e.g. "AdmissionFairSharingAnchorAtQuotaReservation requires AdmissionFairSharing"
# Fix: correct the featureGates list; do NOT blind-restart.

# SIGNATURE 2: webhook panics — the v0.20 webhook is stricter
kubectl -n kueue-system logs deploy/kueue-controller-manager | grep -i "panic\|webhook"
# New rejections at admission: invalid PodSet metadata (WorkloadValidationForPodSetMetadata),
# negative resource requests (WorkloadValidateResourcesAreNonNegative),
# TAS slice-size validation (PodSetSliceSize >= 1).

# SIGNATURE 3: stuck-pending workloads that used to run
kubectl get workloads -A -o json \
  | jq -r '.items[] | select(.status.conditions[]? | select(.type=="QuotaReserved" and .status=="False")) | .metadata.name'
# Check .status.conditions[].reason — if it names node selectors/taints you
# never configured, TASNodeFeasibilityForAllLevels is reclassifying capacity.

The Four Quota Bypasses (Why Your Quotas Were Wrong Before)

Four fixes in v0.20.0 close accounting holes that let workloads consume more than their ClusterQueue quota covered. These are not theoretical: each shipped as a fix with an issue number, and each means your observed utilization pre-upgrade was understated. Expect ClusterQueue usage to look "higher" after the upgrade not because anything changed, but because the ledger finally matches reality.

Bypass fixedMechanism before v0.20Fix in v0.20.0
LeaderWorkerSet group-size scaling (#13279)Raising spec.leaderWorkerTemplate.size on an admitted, Kueue-managed LWS ran more pods per group than reserved quota coveredleaderWorkerTemplate.size immutable while Kueue-managed (LWSImmutableGroupSize, Beta-on). Recreate at new size instead; spec.replicas (group count) stays mutable
Spark under-reservation (#15833)Workloads reserved less CPU/memory than Spark actually requested: cores, memoryOverhead, and Spark's default memory overhead were ignoredEach Pod now reserves its cores as CPU plus at least 384Mi additional memory. Action: raise ClusterQueue quotas for Spark queues on upgrade — every Spark driver/executor got cheaper-to-count, not more expensive to run
Negative-request quota credit (#12838)Negative container resource requests created artificial ClusterQueue quota credit, bypassing configured limitsNegative values floored to zero in accounting, rejected at validation by default (WorkloadValidateResourcesAreNonNegative, Beta-on)
Job parallelism early release (#16111)A Job with parallelism > completions released the quota of still-running Pods once some Pods succeeded, admitting beyond quotaQuota held until Pods terminate, not until siblings succeed

The Spark fix deserves emphasis because it is the one that requires a quota change on upgrade day. Kueue now reserves Spark's cores as CPU and adds the default memory overhead — so a ClusterQueue sized for the old accounting will start rejecting Spark workloads that previously fit. Memory values must use Spark's Java format (512m, 2g); Kubernetes-style 512Mi is rejected. If you run Spark on Kubernetes under Kueue, budget the quota bump before the roll.

The AFS Accounting Anchor: Who Gets Deprioritized Now

For ClusterQueues running UsageBasedAdmissionFairSharing — the admission-time fairness layer, distinct from the older cohort preemption fair sharing — v0.20 moves the accounting anchor. Verified in pkg/scheduler/scheduler.go (only Usage-based ClusterQueues settle penalties; the anchor change applies at quota reservation) and the AFS docs:

Who wins and loses: LocalQueues whose workloads pass AdmissionChecks quickly are unaffected. LocalQueues whose workloads sit in checks for a long time — external provisioning waits, cross-cluster dispatch — now accumulate usage for that waiting time and get deprioritized relative to leaner queues. If a team complains that "our jobs stopped getting in after the upgrade," this anchor move is the first place to look: check kubectl get localqueue <name> -o jsonpath={.status.fairSharing} — the consumedResources and weightedShare fields are the live evidence.

Two open issues temper the fair-sharing story, and honesty requires naming them: #14543 (FairSharing infinite preemption loop, open) is the known failure mode the FairSharingReevaluatePreemptionCandidates alpha fix risks aggravating, and #13320 documents a leaf-CQ non-borrowing starvation case. The v0.20 PrioritizePreemptorWorkloads alpha gate addresses preemption thrash, disabled by default. Fair sharing in Kueue is powerful but still has sharp edges: if you run it, read both issues before the upgrade, and keep the preemption-strategy config (LessThanOrEqualToFinalShare / LessThanInitialShare) pinned in review.

Day-2: Metrics, Alerts, and Blast Radius

The metrics that catch a sick queue

All names below verified against pkg/metrics/metrics.go in the v0.20.0 source (68 registered metrics); the rendered reference is in the Kueue metrics docs. The v0.20 additions worth adding to dashboards:

# Fair-share usage per LocalQueue — watch this shift after the anchor move
kueue_local_queue_admission_fair_sharing_usage

# Unadmitted workloads, by granular reason (UnadmittedWorkloadsObservability, now on)
kueue_unadmitted_workloads

# Pending workloads per ClusterQueue — the classic health signal
kueue_pending_workloads

# NEW v0.20: time from admission to PodsReady recovery, per the recovery flow
kueue_workload_recovery_wait_time_seconds

# NEW v0.20: distinct pending scheduling-equivalence hashes per ClusterQueue
kueue_pending_scheduling_hashes

# NEW v0.20: overlapping preemption target recomputations in-cycle
kueue_preemption_target_recomputations_total

# NEW v0.20: MultiKueue worker-cluster connectivity status
kueue_multikueue_cluster_status

# NEW v0.20: workload total execution time across admissions
kueue_execution_time_seconds

# NEW v0.20: evictions on worker clusters by cluster + reason
kueue_multikueue_workloads_evicted_total

Alerting deltas to configure on upgrade day:

AlertExpression (starting point)Why v0.20 changed it
AFS usage anomalyderiv(kueue_local_queue_admission_fair_sharing_usage[1h]) > 0 sustained on previously-quiet LocalQueuesThe quota-reservation anchor starts charging queues whose workloads wait in AdmissionChecks
Unexpected pending spikekueue_unadmitted_workloads > 0 on queues that ran clean pre-upgradeTASNodeFeasibilityForAllLevels reclassifies domain-fit workloads as inadmissible; Spark accounting raises per-pod reservations
Preemption churnrate(kueue_preempted_workloads_total[15m]) > 0 new baselineRecomputeAssignmentUponPreemptionTargetsOverlap changes preemption timing; #14543's loop risk lives here
Recovery wait regressionhistogram_quantile(0.95, kueue_workload_recovery_wait_time_seconds_bucket) risingNew metric — baseline it now, alert on drift later

Runbook notes

Should You Upgrade This Week?

References & Further Reading