GPU Scheduling for LLM Workloads: When the Queue Is the Product
Sources
- Kueue v0.19.6 release notes (2026-09-24) — quota-leak fixes
- Kueue v0.20.0-rc.2 release notes (2026-09-29) — fair-sharing anchor change
- Kueue docs: Admission Fair Sharing
- Kueue docs: Preemption
- Kueue docs: Dynamic Resource Allocation
- Kueue docs: Workload Priority Class
- Kueue docs: Set Up Dynamic Resource Allocation
- Kueue docs: WaitForPodsReady
- Kubernetes docs: Dynamic Resource Allocation
Every LLM platform hits the same wall, usually around the third week: the GPUs are no longer the constraint — the queue in front of them is. Once more teams want capacity than the fleet can serve, your scheduling layer stops being plumbing and becomes the product. Who waits, who gets preempted, who silently hoards quota while a nightly batch job holds eight H100s at 4% utilization — those decisions are made by whatever admits workloads, and in 2026 that layer on Kubernetes is Kueue.
This is the deep dive into the queue problem half of GPU scheduling: admission ordering, quota charging, preemption, and fairness. The bin-packing half — MIG slicing, time-slicing, DRA prioritized lists, overcommit ratios — lives in our GPU capacity planning guide, and this piece deliberately does not repeat it. Read that one first if you are still deciding how much GPU each workload gets; read this one when the fight is over who gets it first. The device layer that both halves sit on — DRA, GA since Kubernetes 1.34 — is covered in the NVIDIA DRA driver's own docs; production NVIDIA GPU fleets pair that driver with the GPU Operator's DRA setup.
The timing is not arbitrary. Kueue shipped v0.19.6 on September 24, 2026, a patch release whose entire reason to exist is admission correctness: two separate quota-leak bugs that let a fleet admit workloads beyond quota or — worse — strand quota on workloads that no longer exist, starving the queue. Five days later, v0.20.0-rc.2 landed the biggest fair-sharing semantics change since the feature went beta: usage now anchors at quota reservation, not admission. If you run Kueue in front of GPUs, both releases change what your queue actually does.
TL;DR
- The kube-scheduler cannot arbitrate scarcity. Pod priority is per-pod, quota-blind, and has no concept of a team, a budget, or a waiting queue. Admission control is a separate layer, and for job-shaped work it is Kueue (stable line v0.19.6, parallel line v0.18.10, both released 2026-09-24; official support starts at Kubernetes 1.34).
- Priority that matters lives on the Workload, not the Pod.
WorkloadPriorityClassorders the admission queue and drives preemption without touching in-cluster pod priority — so a batch job that preempts nothing at runtime can still win the admission race. - Fairness has two independent knobs. Admission Fair Sharing orders who gets in by LocalQueue usage history; cohort fair sharing plus preemption strategies decides who gets evicted when a queue over-borrows. v0.20 moves the usage anchor to quota reservation, which changes the first knob's behavior even if you change nothing else.
- DRA made GPU quota honest, and complicated. Kueue now charges DRA devices four ways — by count, by extended resource, by MIG counter, by shared capacity — and the v0.19.6 patch fixed a double-charge bug in exactly the path an NVIDIA DRA fleet uses. Pick your accounting mode deliberately; it is a one-way door per DeviceClass.
- Quota leaks are the failure mode nobody dashboards. The two v0.19.6 headline fixes — a StatefulSet scaled to zero that kept its reservation, and a Job with
parallelism > completionsthat handed quota back early — are both silent capacity theft. Watchkueue_pending_workloadsagainst admitted usage, or you will find them by outage.
The Structural Shift: Scarcity Moves the Bottleneck Up the Stack
For fifteen years the Kubernetes scheduling conversation has been about placement: which node, which zone, which topology domain. Kube-scheduler is genuinely excellent at placement. What it has never had is a notion of admission at the fleet level. Its only scarcity signal is pod priority — a number that says "this pod may evict that pod on the same node," which is a node-local argument, not a fleet-level queue discipline.
LLM workloads break the pod-first model in three specific ways:
- The unit of demand is not a pod. A fine-tune is 8 GPUs for 6 hours. An eval sweep is 64 GPUs for 40 minutes. An agent fleet is 200 pods that each want a fraction of a GPU. The thing you need to order, charge, and evict is the job, not the pod — evicting one pod of a distributed training run just wastes the other 63.
- Demand exceeds supply structurally, not transiently. A fleet at 95% GPU utilization is healthy; a cluster at 95% CPU utilization is an incident. The queue is permanent, so its ordering policy is your capacity allocation policy. FIFO-by-default means the loudest submitter wins.
- Cost asymmetry makes wrong admission expensive. Admitting a 6-hour fine-tune ahead of a 40-minute production batch is not an inconvenience; it is four hours of SLA breach. Priority, preemption, and fairness are the levers that make that tradeoff explicit instead of accidental.
Kueue's answer is the Workload: an API object representing the job's total demand (all pod sets, all resources, all DRA claims), which exists in one of two states — pending in a queue, or admitted with a quota reservation. The user's Job (or JobSet, RayJob, MPIJob, plain pod group) is created suspended; Kueue unsuspends it only when the Workload clears the queue. The entire discipline happens before kube-scheduler sees a single pod:
user submits kueue control plane kube cluster
--------------- --------------------------------------------- ----------------
Job (suspended) --> Workload created (podSets, requests, DRA)
|
v
LocalQueue (namespace-scoped router)
|
v
ClusterQueue (quota + policy)
- queueingStrategy: BestEffortFIFO|StrictFIFO
- nominalQuota per flavor, borrowing via Cohort
- preemption: withinClusterQueue / cohort policies
- admissionScope: fair sharing on/off
|
[ scheduler cycle: order by priority,
then fair-share share value ]
|
v
QuotaReserved=True <--- AdmissionChecks
| (ProvisioningRequest,
| external validators)
v
Job unsuspended --> kube-scheduler places pods
| (DRA: allocates the ACTUAL device)
v
PodsReady within timeout?
yes --> running until done
no --> evict + requeue with backoff
(quota returns to the pool)
Two consequences of this shape deserve emphasis before the config walkthrough. First, Kueue admits quota, not devices: in a DRA fleet, Kueue verifies that quota is available and charges it, but it does not know which physical GPU kube-scheduler will allocate — the documented admission-scheduling gap. If anything else grabs the device in between, your workload sits with quota reserved and pods pending. That gap is why WaitForPodsReady exists and why v0.20's new KueueDRADeviceFeasibility gate (alpha, off by default) is worth understanding. Second, queue policy is cluster-admin territory: ClusterQueues are cluster-scoped, so the fairness and preemption decisions in this guide are platform-team decisions, not per-team tuning knobs.
Priority That Doesn't Leak Into the Pod
The first surprise for most teams: Kueue has its own priority object, WorkloadPriorityClass, deliberately decoupled from pod PriorityClass. The split exists because the two priorities answer different questions. Pod priority governs runtime preemption on a node — which is meaningless for a suspended Job with zero pods. Workload priority governs queue ordering and admission preemption — which is exactly where you want your production eval sweep to beat someone's weekend experiment, without giving that eval sweep the right to stomp arbitrary pods at runtime.
apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata:
name: prod-serving
value: 10000
description: "Production inference and eval workloads - win the queue"
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata:
name: batch-experiment
value: 100
description: "Nightly experiments - preemptable by prod"
A job opts in with a label — kueue.x-k8s.io/priority-class: prod-serving — and Kueue copies the value onto the Workload. The label is mutable while the job is suspended, so you can re-triage a stuck workload without deleting it. Three verified behaviors matter operationally:
- A missing class is a hard stop, not a fallback. If the label names a WorkloadPriorityClass that does not exist, Kueue does not create the Workload at all — no silent fallback to pod priority or the default. Since v0.19.6-adjacent releases it emits a
WorkloadPriorityClassNotFoundwarning event on the job, sokubectl describenames the broken reference. Typos here manifest as "my job never starts," which without the event is a support ticket. - Both planes compose. A job can carry a WorkloadPriorityClass for the queue and a pod PriorityClass for runtime; each is used on its own plane. If only a pod PriorityClass is set, Kueue reuses its value for workload priority too — convenient for brownfield fleets, a trap if you did not intend batch jobs to inherit serving pod priority.
- Priorities also order borrowing. Within a cohort, higher-priority workloads get first claim on borrowed quota (controlled by the
PrioritySortingWithinCohortgate). Your priority scheme is thus also your contention-resolution scheme between teams — size the gaps deliberately, not decoratively.
What priority does not give you is protection from a team that simply submits more. A tenant at priority 100 who submits 200 workloads still occupies the whole queue ahead of one priority-90 workload from a quieter tenant. That is what admission fair sharing is for.
Fairness, Take One: Admission Fair Sharing (Who Gets In)
Admission Fair Sharing (beta since v0.15, enabled by default) reorders the admission queue by historical usage of each LocalQueue: workloads from LocalQueues that have consumed less win. It is a scheduler-side soft fairness — no eviction, just ordering — which makes it the safe first fairness knob for a shared LLM fleet where multiple teams target the same ClusterQueue.
The mechanics are worth knowing because they shape tuning. Kueue samples each LocalQueue's usage on an interval and decays it with a configurable half-life; workloads from low-usage queues sort ahead. Each workload that reaches the accounting anchor takes an immediate entry penalty added to its LocalQueue's usage, which is what prevents the submit-100-workloads-before-stats-update exploit. The knobs:
apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
admissionFairSharing:
usageHalfLifeTime: "168h" # usage half-life: 1 week
usageSamplingInterval: "5m" # how often usage is sampled
resourceWeights:
cpu: 1.0
memory: 1.0
nvidia.com/gpu: 8.0 # a GPU-hour costs 8x a CPU-hour
gpu.memory: 8.0 # weight the counter too, if you charge by it
Two production notes. First, enable it per ClusterQueue via admissionScope.admissionMode: UsageBasedAdmissionFairSharing — it is opt-in at the queue level even though the feature gate is on by default. Second, v0.20 changes the accounting anchor: for ClusterQueues using UsageBasedAdmissionFairSharing, usage now counts from quota reservation (when the workload starts holding quota, even while AdmissionChecks are still pending) instead of from admission, behind the AdmissionFairSharingAnchorAtQuotaReservation gate — beta and default-on in v0.20. The practical effect: workloads waiting on slow AdmissionChecks (cluster-autoscaler ProvisioningRequests, for example) no longer sit "free" in fairness accounting while pinning quota. If you explicitly disable the parent AdmissionFairSharing gate, you must disable the anchor gate in the same config or Kueue refuses to start — a deliberately loud failure.
Weight the resources to match what your tenants actually fight over. On an LLM fleet, CPU and memory are noise; set the GPU resource's weight high enough that one GPU-hour of history outweighs a hundred CPU-hours, or your fairness signal will be dominated by CPU-chatty preprocessing jobs.
Fairness, Take Two: Cohorts and Preemption (Who Gets Evicted)
Admission ordering cannot solve the 2 AM version of the problem: the batch team borrowed every idle GPU in the cohort, and production now needs them back. That is preemption, and the policy surface is small but has sharp edges.
A cohort lets ClusterQueues lend each other unused nominal quota. Lending is where the politics live: the borrower is running on resources it does not own, and the owner can take them back. Kueue's two algorithms differ in who pays for that:
| Policy | Trigger | Ordering inside the trigger | Operational character |
|---|---|---|---|
| Classic preemption | Owner above nominal quota again (reclaim), or borrowWithinCohort enabled | Borrowing workloads first, then lowest priority, then most recently admitted | Cheap, predictable, lender-beats-borrower. Fair only in the property sense. |
| Fair sharing preemption | Queue's weighted share exceeds the cohort's — pending queue preempts toward an equal share | Highest-share ClusterQueue's workloads targeted first; strategies decide which workloads qualify | Converges to weighted shares; more eviction churn; needs strategy tuning. |
ClusterQueue-side policy is three fields, and on a GPU fleet each one deserves a deliberate value:
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: prod-gpu
spec:
cohort: gpu-pool
preemption:
withinClusterQueue: LowerPriority # prod preempts prod: priority rules
reclaimWithinCohort: Any # owner reclaims borrowed GPUs at will
borrowWithinCohort:
policy: LowerPriority # when borrowing, only preempt lower-priority
maxPriorityThreshold: 1000 # ...and only below this priority
For fair sharing preemption, the configuration accepts exactly three strategy lists — [LessThanOrEqualToFinalShare], [LessThanInitialShare], or both in that order — and rejects anything else at startup. The difference: LessThanOrEqualToFinalShare will favor evicting smaller workloads to keep the target queue's share math clean (regardless of their priority), while LessThanInitialShare always picks lowest-priority, newest-started workloads in the target. On an LLM fleet I would default to [LessThanInitialShare] alone: it makes preemption targets predictable (your weekend experiments die first), at the cost of slower convergence to perfectly equal shares. Perfect fairness is not worth a fine-tune's checkpoints.
The blast-radius warning that belongs in every runbook: preemption evicts whole Workloads. A preempted 8-GPU fine-tune loses its allocation and requeues; unless the framework checkpoints to durable storage, it restarts from scratch. Preemption is only humane on workloads designed to be interruptible. If your tenants run uncheckpointed training jobs under a preemptable policy, the queue is working as configured and your users are the bug.
DRA Quota: Four Ways to Charge for the Same GPU
This is where GPU scheduling got genuinely harder in 2026, and where the queue's correctness guarantees live. With DRA now the GA device API (stable since Kubernetes 1.34), the old device-plugin integer — nvidia.com/gpu: 2 — is one option among several, and Kueue must charge quota for a device whose size varies by four orders of magnitude. The Kueue DRA integration (beta since v0.18, gates KueueDRAIntegration and friends enabled by default) supports four accounting modes:
| Mode | Charges by | Pod requests via | Status / gate | When to pick it |
|---|---|---|---|---|
| Device count (ResourceClaimTemplate path) | count in the device request | Explicit ResourceClaimTemplate | Beta since v0.18, default-on | Homogeneous whole-GPU claims; simplest correct default |
| Extended resource | The classic resources.requests entry (e.g. nvidia.com/gpu: 1) | Normal resource requests; scheduler auto-creates the ResourceClaim | Beta since v0.19, default-on (KueueDRAIntegrationExtendedResource); needs k8s 1.36+ DRAExtendedResource | Brownfield manifests you cannot rewrite; zero user friction |
| Counter-based (partitionable) | consumesCounters from ResourceSlices (e.g. GPU memory per MIG slice) | ResourceClaimTemplate against a MIG DeviceClass | Beta since v0.19, default-on (KueueDRAIntegrationPartitionableDevices); k8s 1.35+ with DRAPartitionableDevices | MIG fleets — quota in GiB of GPU memory, not "1 device" that might be 10GB or 80GB |
| Capacity-based (shared) | Workload's capacity.requests on the device request | ResourceClaimTemplate with capacity.requests | Alpha in v0.19, default-off (KueueDRAIntegrationConsumableCapacity); k8s 1.36+ with DRAConsumableCapacity | Time-slicing/MPS sharing — variable-capacity claims on one device |
The mode is a property of the deviceClassMappings entry in the Kueue Configuration, so it is a per-DeviceClass one-way door: a DeviceClass either has no sources (count), a counter source, or a capacity source — never both. Configure the mapping, then cover the logical resource name in the ClusterQueue:
# Kueue Configuration — counter-based quota for partitionable (MIG) devices
apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
resources:
deviceClassMappings:
- name: gpu.memory # logical name used in ClusterQueue quotas
deviceClassNames:
- gpu.example.com # your MIG-partitioned DeviceClass
sources:
- counter:
name: memory # must match the driver's ResourceSlice counter key
driver: gpu.example.com
deviceSelector:
cel:
expression: "device.driver == 'gpu.example.com'"
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: gpu-team-a
spec:
cohort: gpu-pool
namespaceSelector: {}
resourceGroups:
- coveredResources: ["gpu.memory"] # quota in GiB, not devices
flavors:
- name: mig-h100
resources:
- name: "gpu.memory"
nominalQuota: 640Gi # 8 x H100 with ~80Gi usable via MIG counters
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
namespace: llm-jobs
name: mig-slice-10gb
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.example.com
---
apiVersion: batch/v1
kind: Job
metadata:
namespace: llm-jobs
name: batch-eval-sweep
labels:
kueue.x-k8s.io/queue-name: team-a-gpu # LocalQueue
kueue.x-k8s.io/priority-class: prod-serving
spec:
suspend: true
template:
spec:
restartPolicy: Never
containers:
- name: eval
image: registry.example.com/eval-runner:latest
resources:
requests:
cpu: "4"
memory: 16Gi
claims:
- name: gpu
resourceClaims:
- name: gpu
resourceClaimTemplateName: mig-slice-10gb
Verified sharp edges, all from the current docs — these bite real fleets:
- Only ExactCount is supported. The
Allallocation mode andFirstAvailableprioritized lists are rejected by the Kueue integration. The prioritized fallback lists that went stable in Kubernetes 1.36 (the "prefer H100, fall back to A100" feature) are not factored into Kueue's quota decisions — use ResourceFlavors for fungibility instead, which is the better tool anyway since flavors carry queueing policy. - Direct ResourceClaims are not supported. Only ResourceClaimTemplate references; a Pod with a direct claim ends up inadmissible, not silently uncharged.
- Counter scales are on you. Kueue does not validate that ClusterQueues sharing a cohort use consistent counter units — a 640Gi queue lending to a 640G queue will produce nonsense accounting. Standardize units per cohort.
- The config is read once at startup. Change
deviceClassMappingson a live cluster and nothing happens until the controller restarts; meanwhile workloads referencing unmapped DeviceClasses are marked inadmissible withDeviceClass <x> is not mapped in DRA configuration, not admitted uncharged. Restart is part of the change procedure, not an accident. - DRA + TopologyAwareScheduling do not compose yet. DRA resources are not accounted in TAS capacity calculations — mixing them yields incorrect topology assignments. Pick one per pool for now.
The v0.19.6 Postmortems: Two Quota Leaks, Zero Alarms
Admission systems fail quietly. Two fixed-in-v0.19.6 bugs are worth internalizing as a class, because they are what quota theft looks like in production:
| Bug (both fixed in v0.19.6) | Mechanism | Why it starves the queue |
|---|---|---|
| StatefulSet scaled to zero kept its reservation (#16128) | An evicted StatefulSet/LeaderWorkerSet workload with no pods never released QuotaReserved; the reservation sat on a phantom | Quota is finite and zero-sum — a phantom holder means real workloads wait for capacity nobody is using. Worse, the queue shows "quota exhausted" while utilization is low. |
parallelism > completions released quota early (#16132) | As soon as some pods succeeded, the Job's remaining pods' quota was returned while they still ran | Opposite direction, same theft: the queue admits new work against quota that is still physically occupied — oversubscription you never configured. |
Neither produces an error event. Both are only visible as a mismatch between reserved and actually running — which is the observability section's whole thesis. The v0.20 line is already tightening this class further: the new UnadmittedWorkloadsObservability default reports granular pending reasons (WaitingForQuota, ExceedsMaxQuota, WaitingForPodsReady, Misconfigured, Suspended) directly in the QuotaReserved condition, and KueueDRADeviceFeasibility (alpha) will stop Kueue from reserving quota for workloads kube-scheduler cannot actually place. Upgrade posture: if you are on v0.18.x, v0.18.10 (same day) carries the equivalents; if you run StatefulSets or partial-admission Jobs under Kueue, v0.19.6 / v0.18.10 is not optional.
Day-2: The Metrics That Catch a Sick Queue
Kueue exposes queue health as first-class Prometheus metrics — the documented set maps one-to-one onto the failure modes above. Minimum viable dashboard, all verified against the metrics reference:
# 1. Admission latency by queue — your real capacity-allocation SLO
histogram_quantile(0.95,
sum by (le, cluster_queue) (rate(kueue_admission_wait_time_seconds_bucket[1h])))
# 2. Queue depth with reason — pending is normal; inadmissible is a config bug
sum by (cluster_queue, status) (kueue_pending_workloads)
# 3. Fair-share drift — a cohort member borrowing far beyond its weight
max by (cluster_queue, cohort) (kueue_cluster_queue_weighted_share)
# 4. Quota-leak canary: reserved-but-not-running (per above, count manually:
# compare quota usage vs pods actually Running per ClusterQueue)
Alerting philosophy: kueue_pending_workloads{status="active"} growing over hours is the system working (demand exceeds supply — by design). status="inadmissible" growing is the system broken — those workloads are not retried until cluster state changes, so a bad flavor name or unmapped DeviceClass strands them forever. Page on the second, trend the first. For bursty agent fleets, pair the queue metrics with the on-demand pending-workloads visibility endpoint, which dumps the current admission-ordering snapshot over an authenticated API — the difference between "300 pending" and "300 pending, all behind one team's 6-hour fine-tune" is the difference between a ticket and an answer.
Finally, the interaction with the serving tier: for inference endpoints, Kueue admission and KEDA-style metric autoscaling are complements, not substitutes — admission governs the batch/training plane (long, interruptible, queue-shaped), autoscaling governs the serving plane (short, latency-shaped, never queued at the gateway). Run both, and keep serving traffic out of the admission queue entirely: a p50 that includes a queue wait is a product decision, not a scheduling one.
A Worked Baseline: The Whole Stack in One Place
Assembled end-to-end for a two-team GPU pool — production serving-adjacent work at priority 10000, experiments preemptable at 100, fair-share ordered at admission, MIG-counter charged, with WaitForPodsReady as the anti-strand safety net:
# 1) Install the stable line (2026-09-24):
# kubectl apply --server-side -f \
# https://github.com/kubernetes-sigs/kueue/releases/download/v0.19.6/manifests.yaml
---
# 2) Priorities: admission ordering only, no runtime pod priority implied
apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata: {name: prod-gpu-work}
value: 10000
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata: {name: experiment-gpu-work}
value: 100
---
# 3) Fairness at admission: usage-history ordering (Configuration)
apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
admissionFairSharing:
usageHalfLifeTime: "168h"
usageSamplingInterval: "5m"
resourceWeights:
cpu: 1.0
memory: 1.0
nvidia.com/gpu: 8.0
waitForPodsReady:
timeout: 30m
blockAdmission: false
requeuingStrategy:
timestamp: Eviction
backoffLimitCount: 5
backoffBaseSeconds: 60
backoffMaxSeconds: 3600
---
# 4) Flavors + cohort: two teams, one GPU pool, borrowing allowed
apiVersion: kueue.x-k8s.io/v1beta2
kind: ResourceFlavor
metadata: {name: mig-h100}
spec:
nodeLabels:
node.kubernetes.io/instance-type: h100-mig
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata: {name: team-prod}
spec:
cohort: gpu-pool
namespaceSelector: {}
queueingStrategy: BestEffortFIFO
admissionScope:
admissionMode: UsageBasedAdmissionFairSharing
preemption:
withinClusterQueue: LowerPriority
reclaimWithinCohort: Any
borrowWithinCohort:
policy: LowerPriority
maxPriorityThreshold: 1000
resourceGroups:
- coveredResources: ["cpu", "memory", "gpu.memory"]
flavors:
- name: mig-h100
resources:
- {name: cpu, nominalQuota: 64}
- {name: memory, nominalQuota: 256Gi}
- {name: gpu.memory, nominalQuota: 320Gi}
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata: {name: team-experiments}
spec:
cohort: gpu-pool
namespaceSelector: {}
queueingStrategy: BestEffortFIFO
preemption:
withinClusterQueue: LowerPriority
reclaimWithinCohort: Any
resourceGroups:
- coveredResources: ["cpu", "memory", "gpu.memory"]
flavors:
- name: mig-h100
resources:
- {name: cpu, nominalQuota: 16}
- {name: memory, nominalQuota: 64Gi}
- {name: gpu.memory, nominalQuota: 160Gi}
Two deliberate choices in that baseline deserve a closing word. The prod queue gets reclaimWithinCohort: Any — when production reclaims borrowed capacity, a mid-flight experiment is evicted and requeued, not queued behind. The experiments queue deliberately omits borrowWithinCohort, so experiments lend idle capacity to prod but only preempt prod work below priority 1000 — which, given the priority ladder, means they effectively never preempt prod. That is the entire fairness architecture in two YAML fields: owners always reclaim, borrowers never displace. Everything else — fair-share ordering at admission, priority within the queue, MIG-counter charging — tunes the throughput and the politics around that one invariant.
References & Further Reading
- Kueue v0.19.6 release notes — the quota-leak fixes (StatefulSet reservation #16128, parallelism-quota early release #16132, DRA overhead accounting #15987), September 24, 2026.
- Kueue v0.20.0-rc.2 release notes — AdmissionFairSharingAnchorAtQuotaReservation, KueueDRADeviceFeasibility, UnadmittedWorkloadsObservability, September 29, 2026.
- Kueue: Admission Fair Sharing and Kueue: Preemption — the fairness and preemption semantics used throughout.
- Kueue: Dynamic Resource Allocation and Set Up DRA — the four quota-accounting modes and their gates, verified against v0.19/v0.20 docs.
- Kubernetes: Dynamic Resource Allocation — upstream concepts; note the canonical URL lives under resource-management, and the scheduling-eviction path redirects.
- Kueue: Workload Priority Class — the admission-priority plane and its label semantics.
- Kueue metrics reference —
kueue_admission_wait_time_seconds,kueue_pending_workloads,kueue_cluster_queue_weighted_share. - KEP-4815: DRA Partitionable Devices, KEP-5075: DRA Consumable Capacity, KEP-5004: DRA Extended Resources — the upstream features behind counter- and capacity-based quota.
- NVIDIA k8s-dra-driver-gpu and the NVIDIA GPU Operator DRA setup — the production driver that publishes the counters and ResourceSlices the counter-based mode charges against.
- Platform Monkey companions: GPU capacity planning for inference fleets (the bin-packing half: MIG arithmetic, overcommit, DRA prioritized lists) and AI fleet architecture patterns (Kueue at fleet scale, budgets as circuit breakers).