GPU Scheduling for LLM Workloads: When the Queue Is the Product

Sources

Every LLM platform hits the same wall, usually around the third week: the GPUs are no longer the constraint — the queue in front of them is. Once more teams want capacity than the fleet can serve, your scheduling layer stops being plumbing and becomes the product. Who waits, who gets preempted, who silently hoards quota while a nightly batch job holds eight H100s at 4% utilization — those decisions are made by whatever admits workloads, and in 2026 that layer on Kubernetes is Kueue.

This is the deep dive into the queue problem half of GPU scheduling: admission ordering, quota charging, preemption, and fairness. The bin-packing half — MIG slicing, time-slicing, DRA prioritized lists, overcommit ratios — lives in our GPU capacity planning guide, and this piece deliberately does not repeat it. Read that one first if you are still deciding how much GPU each workload gets; read this one when the fight is over who gets it first. The device layer that both halves sit on — DRA, GA since Kubernetes 1.34 — is covered in the NVIDIA DRA driver's own docs; production NVIDIA GPU fleets pair that driver with the GPU Operator's DRA setup.

The timing is not arbitrary. Kueue shipped v0.19.6 on September 24, 2026, a patch release whose entire reason to exist is admission correctness: two separate quota-leak bugs that let a fleet admit workloads beyond quota or — worse — strand quota on workloads that no longer exist, starving the queue. Five days later, v0.20.0-rc.2 landed the biggest fair-sharing semantics change since the feature went beta: usage now anchors at quota reservation, not admission. If you run Kueue in front of GPUs, both releases change what your queue actually does.

TL;DR

The Structural Shift: Scarcity Moves the Bottleneck Up the Stack

For fifteen years the Kubernetes scheduling conversation has been about placement: which node, which zone, which topology domain. Kube-scheduler is genuinely excellent at placement. What it has never had is a notion of admission at the fleet level. Its only scarcity signal is pod priority — a number that says "this pod may evict that pod on the same node," which is a node-local argument, not a fleet-level queue discipline.

LLM workloads break the pod-first model in three specific ways:

  1. The unit of demand is not a pod. A fine-tune is 8 GPUs for 6 hours. An eval sweep is 64 GPUs for 40 minutes. An agent fleet is 200 pods that each want a fraction of a GPU. The thing you need to order, charge, and evict is the job, not the pod — evicting one pod of a distributed training run just wastes the other 63.
  2. Demand exceeds supply structurally, not transiently. A fleet at 95% GPU utilization is healthy; a cluster at 95% CPU utilization is an incident. The queue is permanent, so its ordering policy is your capacity allocation policy. FIFO-by-default means the loudest submitter wins.
  3. Cost asymmetry makes wrong admission expensive. Admitting a 6-hour fine-tune ahead of a 40-minute production batch is not an inconvenience; it is four hours of SLA breach. Priority, preemption, and fairness are the levers that make that tradeoff explicit instead of accidental.

Kueue's answer is the Workload: an API object representing the job's total demand (all pod sets, all resources, all DRA claims), which exists in one of two states — pending in a queue, or admitted with a quota reservation. The user's Job (or JobSet, RayJob, MPIJob, plain pod group) is created suspended; Kueue unsuspends it only when the Workload clears the queue. The entire discipline happens before kube-scheduler sees a single pod:

   user submits                 kueue control plane                kube cluster
  ---------------  ---------------------------------------------  ----------------
  Job (suspended) --> Workload created (podSets, requests, DRA)
                          |
                          v
                    LocalQueue (namespace-scoped router)
                          |
                          v
                    ClusterQueue (quota + policy)
                    - queueingStrategy: BestEffortFIFO|StrictFIFO
                    - nominalQuota per flavor, borrowing via Cohort
                    - preemption: withinClusterQueue / cohort policies
                    - admissionScope: fair sharing on/off
                          |
                    [ scheduler cycle: order by priority,
                      then fair-share share value ]
                          |
                          v
                    QuotaReserved=True  <--- AdmissionChecks
                          |                 (ProvisioningRequest,
                          |                  external validators)
                          v
                    Job unsuspended --> kube-scheduler places pods
                          |              (DRA: allocates the ACTUAL device)
                          v
                    PodsReady within timeout?
                      yes --> running until done
                      no  --> evict + requeue with backoff
                                  (quota returns to the pool)

Two consequences of this shape deserve emphasis before the config walkthrough. First, Kueue admits quota, not devices: in a DRA fleet, Kueue verifies that quota is available and charges it, but it does not know which physical GPU kube-scheduler will allocate — the documented admission-scheduling gap. If anything else grabs the device in between, your workload sits with quota reserved and pods pending. That gap is why WaitForPodsReady exists and why v0.20's new KueueDRADeviceFeasibility gate (alpha, off by default) is worth understanding. Second, queue policy is cluster-admin territory: ClusterQueues are cluster-scoped, so the fairness and preemption decisions in this guide are platform-team decisions, not per-team tuning knobs.

Priority That Doesn't Leak Into the Pod

The first surprise for most teams: Kueue has its own priority object, WorkloadPriorityClass, deliberately decoupled from pod PriorityClass. The split exists because the two priorities answer different questions. Pod priority governs runtime preemption on a node — which is meaningless for a suspended Job with zero pods. Workload priority governs queue ordering and admission preemption — which is exactly where you want your production eval sweep to beat someone's weekend experiment, without giving that eval sweep the right to stomp arbitrary pods at runtime.

apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata:
  name: prod-serving
value: 10000
description: "Production inference and eval workloads - win the queue"
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata:
  name: batch-experiment
value: 100
description: "Nightly experiments - preemptable by prod"

A job opts in with a label — kueue.x-k8s.io/priority-class: prod-serving — and Kueue copies the value onto the Workload. The label is mutable while the job is suspended, so you can re-triage a stuck workload without deleting it. Three verified behaviors matter operationally:

What priority does not give you is protection from a team that simply submits more. A tenant at priority 100 who submits 200 workloads still occupies the whole queue ahead of one priority-90 workload from a quieter tenant. That is what admission fair sharing is for.

Fairness, Take One: Admission Fair Sharing (Who Gets In)

Admission Fair Sharing (beta since v0.15, enabled by default) reorders the admission queue by historical usage of each LocalQueue: workloads from LocalQueues that have consumed less win. It is a scheduler-side soft fairness — no eviction, just ordering — which makes it the safe first fairness knob for a shared LLM fleet where multiple teams target the same ClusterQueue.

The mechanics are worth knowing because they shape tuning. Kueue samples each LocalQueue's usage on an interval and decays it with a configurable half-life; workloads from low-usage queues sort ahead. Each workload that reaches the accounting anchor takes an immediate entry penalty added to its LocalQueue's usage, which is what prevents the submit-100-workloads-before-stats-update exploit. The knobs:

apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
admissionFairSharing:
  usageHalfLifeTime: "168h"        # usage half-life: 1 week
  usageSamplingInterval: "5m"      # how often usage is sampled
  resourceWeights:
    cpu: 1.0
    memory: 1.0
    nvidia.com/gpu: 8.0            # a GPU-hour costs 8x a CPU-hour
    gpu.memory: 8.0                # weight the counter too, if you charge by it

Two production notes. First, enable it per ClusterQueue via admissionScope.admissionMode: UsageBasedAdmissionFairSharing — it is opt-in at the queue level even though the feature gate is on by default. Second, v0.20 changes the accounting anchor: for ClusterQueues using UsageBasedAdmissionFairSharing, usage now counts from quota reservation (when the workload starts holding quota, even while AdmissionChecks are still pending) instead of from admission, behind the AdmissionFairSharingAnchorAtQuotaReservation gate — beta and default-on in v0.20. The practical effect: workloads waiting on slow AdmissionChecks (cluster-autoscaler ProvisioningRequests, for example) no longer sit "free" in fairness accounting while pinning quota. If you explicitly disable the parent AdmissionFairSharing gate, you must disable the anchor gate in the same config or Kueue refuses to start — a deliberately loud failure.

Weight the resources to match what your tenants actually fight over. On an LLM fleet, CPU and memory are noise; set the GPU resource's weight high enough that one GPU-hour of history outweighs a hundred CPU-hours, or your fairness signal will be dominated by CPU-chatty preprocessing jobs.

Fairness, Take Two: Cohorts and Preemption (Who Gets Evicted)

Admission ordering cannot solve the 2 AM version of the problem: the batch team borrowed every idle GPU in the cohort, and production now needs them back. That is preemption, and the policy surface is small but has sharp edges.

A cohort lets ClusterQueues lend each other unused nominal quota. Lending is where the politics live: the borrower is running on resources it does not own, and the owner can take them back. Kueue's two algorithms differ in who pays for that:

PolicyTriggerOrdering inside the triggerOperational character
Classic preemptionOwner above nominal quota again (reclaim), or borrowWithinCohort enabledBorrowing workloads first, then lowest priority, then most recently admittedCheap, predictable, lender-beats-borrower. Fair only in the property sense.
Fair sharing preemptionQueue's weighted share exceeds the cohort's — pending queue preempts toward an equal shareHighest-share ClusterQueue's workloads targeted first; strategies decide which workloads qualifyConverges to weighted shares; more eviction churn; needs strategy tuning.

ClusterQueue-side policy is three fields, and on a GPU fleet each one deserves a deliberate value:

apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
  name: prod-gpu
spec:
  cohort: gpu-pool
  preemption:
    withinClusterQueue: LowerPriority      # prod preempts prod: priority rules
    reclaimWithinCohort: Any               # owner reclaims borrowed GPUs at will
    borrowWithinCohort:
      policy: LowerPriority                # when borrowing, only preempt lower-priority
      maxPriorityThreshold: 1000          # ...and only below this priority

For fair sharing preemption, the configuration accepts exactly three strategy lists — [LessThanOrEqualToFinalShare], [LessThanInitialShare], or both in that order — and rejects anything else at startup. The difference: LessThanOrEqualToFinalShare will favor evicting smaller workloads to keep the target queue's share math clean (regardless of their priority), while LessThanInitialShare always picks lowest-priority, newest-started workloads in the target. On an LLM fleet I would default to [LessThanInitialShare] alone: it makes preemption targets predictable (your weekend experiments die first), at the cost of slower convergence to perfectly equal shares. Perfect fairness is not worth a fine-tune's checkpoints.

The blast-radius warning that belongs in every runbook: preemption evicts whole Workloads. A preempted 8-GPU fine-tune loses its allocation and requeues; unless the framework checkpoints to durable storage, it restarts from scratch. Preemption is only humane on workloads designed to be interruptible. If your tenants run uncheckpointed training jobs under a preemptable policy, the queue is working as configured and your users are the bug.

DRA Quota: Four Ways to Charge for the Same GPU

This is where GPU scheduling got genuinely harder in 2026, and where the queue's correctness guarantees live. With DRA now the GA device API (stable since Kubernetes 1.34), the old device-plugin integer — nvidia.com/gpu: 2 — is one option among several, and Kueue must charge quota for a device whose size varies by four orders of magnitude. The Kueue DRA integration (beta since v0.18, gates KueueDRAIntegration and friends enabled by default) supports four accounting modes:

ModeCharges byPod requests viaStatus / gateWhen to pick it
Device count (ResourceClaimTemplate path)count in the device requestExplicit ResourceClaimTemplateBeta since v0.18, default-onHomogeneous whole-GPU claims; simplest correct default
Extended resourceThe classic resources.requests entry (e.g. nvidia.com/gpu: 1)Normal resource requests; scheduler auto-creates the ResourceClaimBeta since v0.19, default-on (KueueDRAIntegrationExtendedResource); needs k8s 1.36+ DRAExtendedResourceBrownfield manifests you cannot rewrite; zero user friction
Counter-based (partitionable)consumesCounters from ResourceSlices (e.g. GPU memory per MIG slice)ResourceClaimTemplate against a MIG DeviceClassBeta since v0.19, default-on (KueueDRAIntegrationPartitionableDevices); k8s 1.35+ with DRAPartitionableDevicesMIG fleets — quota in GiB of GPU memory, not "1 device" that might be 10GB or 80GB
Capacity-based (shared)Workload's capacity.requests on the device requestResourceClaimTemplate with capacity.requestsAlpha in v0.19, default-off (KueueDRAIntegrationConsumableCapacity); k8s 1.36+ with DRAConsumableCapacityTime-slicing/MPS sharing — variable-capacity claims on one device

The mode is a property of the deviceClassMappings entry in the Kueue Configuration, so it is a per-DeviceClass one-way door: a DeviceClass either has no sources (count), a counter source, or a capacity source — never both. Configure the mapping, then cover the logical resource name in the ClusterQueue:

# Kueue Configuration — counter-based quota for partitionable (MIG) devices
apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
resources:
  deviceClassMappings:
  - name: gpu.memory                 # logical name used in ClusterQueue quotas
    deviceClassNames:
    - gpu.example.com                # your MIG-partitioned DeviceClass
    sources:
    - counter:
        name: memory                 # must match the driver's ResourceSlice counter key
        driver: gpu.example.com
        deviceSelector:
          cel:
            expression: "device.driver == 'gpu.example.com'"
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
  name: gpu-team-a
spec:
  cohort: gpu-pool
  namespaceSelector: {}
  resourceGroups:
  - coveredResources: ["gpu.memory"]   # quota in GiB, not devices
    flavors:
    - name: mig-h100
      resources:
      - name: "gpu.memory"
        nominalQuota: 640Gi            # 8 x H100 with ~80Gi usable via MIG counters
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  namespace: llm-jobs
  name: mig-slice-10gb
spec:
  spec:
    devices:
      requests:
      - name: gpu
        exactly:
          deviceClassName: gpu.example.com
---
apiVersion: batch/v1
kind: Job
metadata:
  namespace: llm-jobs
  name: batch-eval-sweep
  labels:
    kueue.x-k8s.io/queue-name: team-a-gpu   # LocalQueue
    kueue.x-k8s.io/priority-class: prod-serving
spec:
  suspend: true
  template:
    spec:
      restartPolicy: Never
      containers:
      - name: eval
        image: registry.example.com/eval-runner:latest
        resources:
          requests:
            cpu: "4"
            memory: 16Gi
          claims:
          - name: gpu
      resourceClaims:
      - name: gpu
        resourceClaimTemplateName: mig-slice-10gb

Verified sharp edges, all from the current docs — these bite real fleets:

The v0.19.6 Postmortems: Two Quota Leaks, Zero Alarms

Admission systems fail quietly. Two fixed-in-v0.19.6 bugs are worth internalizing as a class, because they are what quota theft looks like in production:

Bug (both fixed in v0.19.6)MechanismWhy it starves the queue
StatefulSet scaled to zero kept its reservation (#16128)An evicted StatefulSet/LeaderWorkerSet workload with no pods never released QuotaReserved; the reservation sat on a phantomQuota is finite and zero-sum — a phantom holder means real workloads wait for capacity nobody is using. Worse, the queue shows "quota exhausted" while utilization is low.
parallelism > completions released quota early (#16132)As soon as some pods succeeded, the Job's remaining pods' quota was returned while they still ranOpposite direction, same theft: the queue admits new work against quota that is still physically occupied — oversubscription you never configured.

Neither produces an error event. Both are only visible as a mismatch between reserved and actually running — which is the observability section's whole thesis. The v0.20 line is already tightening this class further: the new UnadmittedWorkloadsObservability default reports granular pending reasons (WaitingForQuota, ExceedsMaxQuota, WaitingForPodsReady, Misconfigured, Suspended) directly in the QuotaReserved condition, and KueueDRADeviceFeasibility (alpha) will stop Kueue from reserving quota for workloads kube-scheduler cannot actually place. Upgrade posture: if you are on v0.18.x, v0.18.10 (same day) carries the equivalents; if you run StatefulSets or partial-admission Jobs under Kueue, v0.19.6 / v0.18.10 is not optional.

Day-2: The Metrics That Catch a Sick Queue

Kueue exposes queue health as first-class Prometheus metrics — the documented set maps one-to-one onto the failure modes above. Minimum viable dashboard, all verified against the metrics reference:

# 1. Admission latency by queue — your real capacity-allocation SLO
histogram_quantile(0.95,
  sum by (le, cluster_queue) (rate(kueue_admission_wait_time_seconds_bucket[1h])))

# 2. Queue depth with reason — pending is normal; inadmissible is a config bug
sum by (cluster_queue, status) (kueue_pending_workloads)

# 3. Fair-share drift — a cohort member borrowing far beyond its weight
max by (cluster_queue, cohort) (kueue_cluster_queue_weighted_share)

# 4. Quota-leak canary: reserved-but-not-running (per above, count manually:
#    compare quota usage vs pods actually Running per ClusterQueue)

Alerting philosophy: kueue_pending_workloads{status="active"} growing over hours is the system working (demand exceeds supply — by design). status="inadmissible" growing is the system broken — those workloads are not retried until cluster state changes, so a bad flavor name or unmapped DeviceClass strands them forever. Page on the second, trend the first. For bursty agent fleets, pair the queue metrics with the on-demand pending-workloads visibility endpoint, which dumps the current admission-ordering snapshot over an authenticated API — the difference between "300 pending" and "300 pending, all behind one team's 6-hour fine-tune" is the difference between a ticket and an answer.

Finally, the interaction with the serving tier: for inference endpoints, Kueue admission and KEDA-style metric autoscaling are complements, not substitutes — admission governs the batch/training plane (long, interruptible, queue-shaped), autoscaling governs the serving plane (short, latency-shaped, never queued at the gateway). Run both, and keep serving traffic out of the admission queue entirely: a p50 that includes a queue wait is a product decision, not a scheduling one.

A Worked Baseline: The Whole Stack in One Place

Assembled end-to-end for a two-team GPU pool — production serving-adjacent work at priority 10000, experiments preemptable at 100, fair-share ordered at admission, MIG-counter charged, with WaitForPodsReady as the anti-strand safety net:

# 1) Install the stable line (2026-09-24):
#    kubectl apply --server-side -f \
#      https://github.com/kubernetes-sigs/kueue/releases/download/v0.19.6/manifests.yaml
---
# 2) Priorities: admission ordering only, no runtime pod priority implied
apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata: {name: prod-gpu-work}
value: 10000
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: WorkloadPriorityClass
metadata: {name: experiment-gpu-work}
value: 100
---
# 3) Fairness at admission: usage-history ordering (Configuration)
apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
admissionFairSharing:
  usageHalfLifeTime: "168h"
  usageSamplingInterval: "5m"
  resourceWeights:
    cpu: 1.0
    memory: 1.0
    nvidia.com/gpu: 8.0
waitForPodsReady:
  timeout: 30m
  blockAdmission: false
  requeuingStrategy:
    timestamp: Eviction
    backoffLimitCount: 5
    backoffBaseSeconds: 60
    backoffMaxSeconds: 3600
---
# 4) Flavors + cohort: two teams, one GPU pool, borrowing allowed
apiVersion: kueue.x-k8s.io/v1beta2
kind: ResourceFlavor
metadata: {name: mig-h100}
spec:
  nodeLabels:
    node.kubernetes.io/instance-type: h100-mig
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata: {name: team-prod}
spec:
  cohort: gpu-pool
  namespaceSelector: {}
  queueingStrategy: BestEffortFIFO
  admissionScope:
    admissionMode: UsageBasedAdmissionFairSharing
  preemption:
    withinClusterQueue: LowerPriority
    reclaimWithinCohort: Any
    borrowWithinCohort:
      policy: LowerPriority
      maxPriorityThreshold: 1000
  resourceGroups:
  - coveredResources: ["cpu", "memory", "gpu.memory"]
    flavors:
    - name: mig-h100
      resources:
      - {name: cpu, nominalQuota: 64}
      - {name: memory, nominalQuota: 256Gi}
      - {name: gpu.memory, nominalQuota: 320Gi}
---
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata: {name: team-experiments}
spec:
  cohort: gpu-pool
  namespaceSelector: {}
  queueingStrategy: BestEffortFIFO
  preemption:
    withinClusterQueue: LowerPriority
    reclaimWithinCohort: Any
  resourceGroups:
  - coveredResources: ["cpu", "memory", "gpu.memory"]
    flavors:
    - name: mig-h100
      resources:
      - {name: cpu, nominalQuota: 16}
      - {name: memory, nominalQuota: 64Gi}
      - {name: gpu.memory, nominalQuota: 160Gi}

Two deliberate choices in that baseline deserve a closing word. The prod queue gets reclaimWithinCohort: Any — when production reclaims borrowed capacity, a mid-flight experiment is evicted and requeued, not queued behind. The experiments queue deliberately omits borrowWithinCohort, so experiments lend idle capacity to prod but only preempt prod work below priority 1000 — which, given the priority ladder, means they effectively never preempt prod. That is the entire fairness architecture in two YAML fields: owners always reclaim, borrowers never displace. Everything else — fair-share ordering at admission, priority within the queue, MIG-counter charging — tunes the throughput and the politics around that one invariant.

References & Further Reading