GPU Capacity Planning for Inference Fleets: Bin-Packing H100s Without the 2 AM Page
Sources
- Kubernetes blog: v1.36 DRA updates (prioritized list stable)
- Kubernetes blog: v1.34 — DRA has graduated to GA
- Kubernetes docs: Dynamic Resource Allocation concepts
- NVIDIA GPU Operator docs: DRA Driver for GPUs
- NVIDIA DRA driver for GPUs (SIG docs)
- Cloudflare blog: How Cloudflare runs more AI models on fewer GPUs (Omni)
- Google Cloud: Best practices for autoscaling LLM inference on GKE
- Character.AI: Optimizing AI inference (Part Deux)
- Character.AI: Optimizing AI inference (original post)
- AWS News Blog: up to 45% price reduction for EC2 NVIDIA GPU instances (Jun 2025)
- Amazon EC2 P5 instances
- Lambda GPU cloud pricing
GPU capacity planning for LLM inference is broken in most orgs, and the failure looks identical everywhere: a fleet of H100s provisioned for peak traffic sits at 5–15% utilization while a queue somewhere quietly breaches its SLO. The instinctive fix — buy more GPUs or shard harder — treats the symptom. The real problem is that capacity planning for inference is a memory problem disguised as a compute problem, and almost every tool in your stack treats it as the opposite.
This guide is the planning playbook that falls out of what production fleets actually do in 2026: Cloudflare packing 13 models onto a single GPU at ~400% memory over-commit, Character.AI redesigning attention before buying more hardware, and Kubernetes graduating DRA's prioritized device lists to stable in 1.36 so a claim can say “give me an H100, fall back to an A100 if none are free.” The order of operations matters: architecture first, partitioning second, elastic fleet last. Doing it in reverse is how you end up with a $200k/month fleet serving a product that never needed it.
TL;DR
- Capacity for inference is KV-cache memory, not FLOPs. A GPU is “full” when its KV pool can’t hold more concurrent conversations, long before its compute saturates. Plan in tokens-resident, not tokens/second.
- DRA finally makes GPUs a schedulable resource. The device-plugin integer model is gone: DRA went GA in 1.34, the gate locked on in 1.35, and 1.36 made prioritized fallback lists stable — ordered preference (“H100, else A100”) instead of one hardcoded device model.
- Time-slicing and MIG are the bin-packing tools. The NVIDIA DRA driver exposes GPU sharing (time-slicing, MPS) and MIG slicing directly through DeviceClasses, so a 10GB inference pod can claim 1/7th of an H100 instead of a whole card.
- Over-commit is a real technique with a real price. Cloudflare’s Omni runs 13 models per GPU by intercepting CUDA allocations into unified memory — and pays ~156ms cold-swap latency when a swapped-out model is called. Over-commit long-tail models; never over-commit your flagship.
- Scale on queue size, not GPU utilization. Google’s own GKE guidance ranks queue size and batch size above GPU utilization because utilization doesn’t map to latency. A static 16-GPU fleet costs $1,532/day at Lambda rates; queue-driven elasticity at the same SLO costs $958/day — a $17,237/month difference (Decimal-verified, 2026-09-28 rates).
- When NOT to pack: latency-critical interactive flagship serving, training jobs, and any fleet under 8 GPUs — the orchestration tax exceeds the savings. Numbers below.
The problem: GPUs are scheduled as integers, capacity is a memory curve
The Kubernetes device-plugin model that most fleets still run treats a GPU as an opaque countable integer: a pod either gets nvidia.com/gpu: 1 or nothing. That model has no way to express sharing, per-workload configuration, capability constraints, or topology requirements. Meanwhile the workload’s actual constraint is a memory curve: how many tokens the KV cache can hold for the model’s context length, which sets the concurrency ceiling of the card.
Run the numbers for a single H100 80GB serving a 32B dense GQA model in bf16. We computed the per-token KV cost and pool size for our vLLM review with Decimal-verified math; the shape of the result is what matters here:
Per-token KV bytes (dense GQA, 32B-class, 64 layers):
8 KV heads x 128 head_dim x 64 layers x 2 (K+V) x 2 bytes
= 262,144 bytes/token = 256 KiB/token
Usable H100 80GB memory pool at utilization 0.92:
80 GiB x 0.92 ~= 73.7 GiB
- model weights (32B x 2 bytes) ~64.0 GiB * does not fit; use fp8/AWQ
- 8B-class model instead (16 GiB) ~16.0 GiB
---------------------------------------
KV pool for 8B model ~57.7 GiB
Concurrent conversations (8K context, 8K gen):
57.7 GiB / (16,384 tokens x 256 KiB) ~= 14 concurrent streams
Three things fall out of that table, and they are the whole ballgame for capacity planning:
- Concurrency is
KV_pool / (context × bytes_per_token), full stop. An 8B model on one H100 holds roughly 14 concurrent 16K-token conversations in bf16 with default preallocation. That is the real capacity of a $4–7/hour card, and it has nothing to do with FLOPs. - Quantization is the first capacity lever, not the last. Cutting KV bytes (fp8 KV cache, GQA head pruning, or MLA architectures) multiplies concurrency at zero hardware cost. Character.AI’s inference stack (int8 attention with cross-layer KV sharing and hybrid attention horizons) exists precisely because “the KV cache is the bottleneck” — they redesigned attention before scaling the fleet.
- GPU utilization is the wrong signal at every level. GKE’s own autoscaling guidance says
DCGM_FI_DEV_GPU_UTIL“does not measure how much work is being done while the GPU is active,” making it hard to map to latency;DCGM_FI_DEV_FB_USEDwon’t scale down preallocating servers like vLLM/TGI. More on the right signal below.
The planning unit that matters is tokens-resident: how many concurrent request-streams the fleet can hold at your p95 context length. Once you internalize that, every packing decision in this guide is arithmetic.
The 2026 packing toolkit: what changed and what it costs
For five years the answer to “can two workloads share a GPU?” was “time-slicing via the device plugin, and may God have mercy on your tail latency.” That changed. The current toolkit, in the order you should reach for it:
1. DRA: GPUs as claims, not integers
Dynamic Resource Allocation graduated to GA in Kubernetes 1.34 (September 2025), and the DynamicResourceAllocation feature gate is part of the default enabled set since 1.34/1.35 — on 1.32/1.33 you must enable it manually, per the NVIDIA DRA driver docs. The model is closer to PersistentVolumes than to device plugins: a workload declares what it needs in a ResourceClaim (possibly via a ResourceClaimTemplate), the driver publishes available devices as ResourceSlices, and a DeviceClass defines how a class of devices may be allocated — including sharing configuration that a device plugin could never express.
The graduation that matters for capacity planning landed in 1.36: prioritized lists went stable. From the official 1.36 DRA blog: instead of hardcoding a request for a specific device model, a claim can specify an ordered list of preferences — “give me an H100, but if none are available, fall back to an A100.” The scheduler evaluates preferences in order, which the release notes describe as drastically improving scheduling flexibility and cluster utilization. For a capacity planner, prioritized lists convert device heterogeneity from an ops headache into a bin-packing primitive: your premium tier claims H100s first and degrades gracefully; your batch tier claims A100s and only lands on H100s when they’d otherwise idle.
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: inference-gpu-preferred
spec:
spec:
resourceClassName: gpu.nvidia.com
devices:
requests:
- name: gpu
exactly:
# first preference: 80GB Hopper with 700+ GB/s memory bandwidth
- deviceClassName: gpu.nvidia.com
selectors:
- gpu.nvidia.com/model: H100
gpu.nvidia.com/memory: "80"
gpu.nvidia.com/brand: NVIDIA-H100-80GB
# fallback: any 80GB card (A100 80GB acceptable)
- deviceClassName: gpu.nvidia.com
selectors:
- gpu.nvidia.com/memory: "80"
Two operational caveats before you lean on this. First, driver maturity: the NVIDIA GPU Operator DRA driver lists some DRA features with alpha support status, requires Kubernetes v1.34.2+, and a cluster can run either the DRA-mode GPUCluster resource or the device-plugin ClusterPolicy — but not both. That is a real migration step, not a flag flip. Second, quota correctness: Kueue’s GPU admission control is still catching up with sharing — Kueue issue #15550 tracks quota-correct admission for DRA DeviceClasses with time-slicing/MPS sharing config. Until that lands, a Kueue quota counts claims, not slice-hours, and your chargeback math will disagree with your scheduler’s.
2. Time-slicing and MPS: the utilitarian packing tools
Below full-GPU claims, the NVIDIA DRA driver lets a single GPU allocation be shared across containers with time-slicing or Multi-Process Service (MPS), and lets one GPU be shared across independent claims via DRA “consumable capacity.” Time-slicing round-robins the GPU between contexts with no memory isolation; MPS gives concurrent kernel execution with a shared address space. Both raise packing density on underutilized cards; both blur your latency floor. The planning rule: time-slice your long-tail models (internal tools, batch enrichment, low-QPS endpoints) and keep flagship interactive serving on exclusive devices.
3. MIG: hardware partitioning when isolation is non-negotiable
MIG physically partitions an H100 into up to seven isolated instances with independent memory bandwidth and fault domains — a 1g.10gb slice is 10GB of memory and a guaranteed fraction of SMs. If one tenant’s runaway request can’t be allowed to OOM the neighbor’s inference pod, MIG is the only sharing mode that gives you that guarantee in hardware. The DRA driver manages MIG slices as first-class claimable devices, so a MIG slice can carry its own DeviceClass with per-class sharing config. The cost is flexibility: MIG partitioning is relatively static (a re-partition drains the card), so most fleets partition once at fleet-provisioning time based on the tenant mix they actually have.
H100 80GB, MIG profile 1g.10gb (1 GPU slice, 10GB GDDR):
max 7x 1g.10gb slices per card (70GB usable; ~10GB reserved)
Tenant mix that motivated the partitioning:
- 4x internal chatbot endpoints (8B models, 4x 10GB slices)
- 2x embedding/rerank services (2x 10GB slices)
- 1x canary/golden-path slot (1x 10GB slice, kept free)
Fleet math:
1 H100 -> 7 inference slots
14 H100 fleet -> 98 slots -> serves ~40 internal models
without MIG: same models need 40 GPUs (1 model/GPU)
That example is the entire MIG pitch: the same 14 cards serve 40 models instead of 14 at the cost of a one-time partitioning decision and a hard per-slice memory ceiling. The failure mode is equally simple — if a tenant’s model grows past 10GB, its slice OOMs at the partition boundary and no scheduler trick recovers it. MIG is a bet on a stable tenant mix.
4. Over-commit: Cloudflare’s Omni, and the latency you pay
The most aggressive packing data point in public is Cloudflare’s. Their Omni platform for Workers AI packs many small, low-volume models onto edge GPUs that would otherwise idle. The published mechanics: per-model process isolation with per-model Python virtual environments managed by uv, a CUDA stub library that intercepts cudaMalloc/cuMalloc calls and forces allocations into unified memory, and an overridden cudaMemGetInfo so each model only sees a subset of GPU memory. The result they publish: 13 models on a single GPU at ~400% memory over-commit, saving up to 4 GPUs per node in their deployment.
Over-commit works because inactive models migrate to CPU memory while active models stay hot on the GPU. The bill is a cold-start swap when a swapped-out model is called: moving a model with ~5GiB of weights and caches back onto the GPU takes ~156ms at PCIe 4.0’s ~32 GB/s. For Workers AI’s workload — thousands of small, low-QPS models where a sub-200ms cold-path is fine — that trade is obviously correct. For your flagship p50-sensitive endpoint, it is obviously wrong. Over-commit is a long-tail tool. Note also what Omni is not: it is not Kubernetes DRA, it is Cloudflare’s bespoke control plane — a reminder that the biggest packing wins still live below the orchestrator, in the runtime layer.
The fleet pattern: static base + elastic top, scheduled by queue depth
Packing decides how many models fit on a card. Fleet sizing decides how many cards you own. The pattern that production fleets converge on is a two-tier fleet:
+----------------------------+
static base | N GPUs, always on | covers p50 demand at
(MIG-partitioned, | MIG slices per tenant | 24/7 duty cycle
time-sliced) +----------------------------+
^
| overflow / peak
v
+----------------------------+
elastic top | 0..M GPUs, scale on | covers p95-p99 peaks,
(exclusive GPUs, | queue depth (HPA + | drains to zero off-peak
DRA claims) | KEDA on server metrics) |
The static tier is where packing (MIG, time-slicing, prioritized claims) pays: it runs at high utilization because it is sized for the floor, not the peak. The elastic tier is where scheduling pays: it exists only to absorb peaks, and it should be exclusive-GPU (no packing) because peak traffic is exactly when you want zero contention. What triggers the elastic tier is the difference between a fleet that saves money and one that just adds a second idle fleet.
The right signal: queue size, not GPU utilization
Google’s GKE autoscaling best practices for LLM inference is the most direct public guidance on this, and it ranks the signals:
| Signal | What it measures | Google’s guidance |
|---|---|---|
| Queue size (server metric) | Requests awaiting processing in the server queue | Use to maximize throughput and cost within a target latency threshold; resilient to traffic fluctuation |
| Batch size (server metric) | Requests currently undergoing inference | Use to reach lower latency thresholds than queue size can |
GPU utilization (DCGM_FI_DEV_GPU_UTIL) | Duty cycle of GPU activity | Does not measure work done while active; hard to map to latency |
GPU memory (DCGM_FI_DEV_FB_USED) | Frame-buffer used at a point in time | Only scales UP for preallocating servers (vLLM, TGI); never scales down |
| CPU/memory | Generic utilization | Not recommended as sole indicators for GPU inference |
The queue-size formula from the same doc: pick your target latency threshold, find the maximum throughput that holds it, then set the HPA target queue size to the level your fleet sustains at that throughput. And be mindful of the HPA tolerance, a default 0.1 no-action band around the target that dampens oscillation — set your target with that band in mind or your fleet will sawtooth. vLLM and TGI expose the queue and batch metrics natively; vLLM’s Prometheus surface is documented in its metrics design doc and we cover the operational details in our vLLM review.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: inference-elastic-top
spec:
scaleTargetRef:
name: vllm-elastic
minReplicaCount: 0 # elastic tier drains to zero off-peak
maxReplicaCount: 8
cooldownPeriod: 600 # 10 min: inference nodes drain slowly (weights load)
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring:9090
metricName: vllm:num_requests_waiting
# threshold: requests waiting per replica before scale-out
threshold: "4"
query: avg(vllm:num_requests_waiting{deployment="vllm-elastic"})
Two production notes on that manifest. First, minReplicaCount: 0 plus cooldownPeriod: 600 encodes the real cost of scaling inference: a replica that scales up must load model weights (tens of seconds to minutes for large models), so you scale early on queue depth, not late. Second, the KEDA threshold is per-trigger; divide your fleet-wide queue target by replica count and re-derive it whenever the base tier’s packing changes.
The cost math: what packing and elasticity actually save (2026-09-28 rates)
Every number in this section is Decimal-verified against published rates on 2026-09-28 and rounded to cents. The input rates:
| Provider / SKU | Config | On-demand rate | Per-GPU |
|---|---|---|---|
| AWS p5.48xlarge | 8× H100 (SXM5, NVLink) | $55.04/hr | $6.88/hr |
| GCP a3-highgpu-8g | 8× H100 | $87.8325/hr | $10.98/hr |
| Lambda 1× H100 SXM 80GB | 1 GPU, 208 vCPU | $3.99/hr | $3.99/hr |
Context for the AWS number: AWS cut P5 on-demand prices 44% in June 2025 (from the May 31, 2025 baseline, per the announcement blog). Context for the spread: hyperscaler H100s cost 1.7–2.8× a specialist GPU cloud at on-demand rates, and that ratio is the entire economic argument for the two-tier fleet pattern — the elastic tier’s cost depends on who you can drain to zero fastest.
What 40% idle actually costs
64x H100 fleet, 40% of GPU-hours idle:
AWS P5 per-GPU ($6.88/hr): $4,227.07/day $126,812.16/month
GCP A3 per-GPU ($10.98/hr): $6,745.54/day $202,366.08/month
Lambda per-GPU ($3.99/hr): $2,451.46/day $73,543.68/month
Same fleet, 20% idle (packing + elastic tier working):
Lambda per-GPU: $1,225.73/day $36,771.84/month
Read that as: the difference between a packed-and-elastic fleet and a static one is roughly $36,772/month at Lambda rates for a 64-GPU fleet — and roughly $100k/month if the same fleet lives on GCP A3 on-demand. Packing is not a tuning exercise; at fleet scale it is a headcount-sized line item.
Static vs. elastic, same SLO
Static: 16 GPUs x $3.99 x 24h = $1,532.16/day
Elastic: 10 avg GPUs x $3.99 x 24h = $957.60/day
----------
$ 574.56/day saved
$17,236.80/month saved
The elastic fleet in that example averages 10 GPUs against a peak of 16 — the average comes from your traffic’s duty cycle, and the entire saving depends on the elastic tier draining to zero off-peak. If your traffic is flat around the clock (internal developer tools often are), elasticity saves nothing and the static packed fleet is the honest answer.
The interconnect tax on multi-node serving
When a model spans nodes (weights > single-card memory, or TP/EP sharding), interconnect bandwidth becomes a capacity-planning input, not a footnote. Time to move one node’s shard of a 70GB bf16 weight set (Decimal-verified):
PCIe 4.0 x16 ~32 GB/s 70 GB / 32 GB/s = 2.188 s
NVLink (H100) ~900 GB/s 70 GB / 900 GB/s = 0.078 s
400G IB ~50 GB/s 70 GB / 50 GB/s = 1.400 s
That 28× gap between PCIe and NVLink is why p5/a3 instances bundle NVSwitch and why “just shard across cheap single-GPU nodes” is usually false economy for models above ~40GB of weights. It is also why Cloudflare’s 156ms cold-swap number is tolerable at their scale (small models, PCIe-class edge hardware, latency-insensitive long-tail) and would be catastrophic on a flagship interactive endpoint.
Real deployments: who runs what
| Deployment | Pattern in production | Published evidence |
|---|---|---|
| Cloudflare Workers AI (Omni) | Per-model process isolation + CUDA-stub memory over-commit; 13 models/GPU at ~400% over-commit; cold-swap ~156ms for a 5GiB model over PCIe 4.0 | Cloudflare technical deep-dive |
| Character.AI | Architecture-first: int8 attention, cross-layer KV sharing, hybrid attention horizons — attack the KV bottleneck before scaling fleet size | Character.AI part deux |
| Google (GKE guidance) | Queue-size/batch-size HPA over GPU-utilization autoscaling; tolerance 0.1; DCGM metrics only for scale-up | GKE autoscaling best practices |
| Kubernetes + NVIDIA DRA | ResourceClaim-first allocation with prioritized fallback (stable in 1.36); MIG/time-slicing/MPS as DeviceClass-configured sharing | k8s 1.36 DRA blog, NVIDIA DRA driver docs |
A pattern across all of them: nobody scaled their way out of a KV bottleneck. Cloudflare and Character.AI both attacked memory first, then packed, then bought hardware. The fleets that struggle are the ones that buy first.
The playbook, in order
- Compute your tokens-resident capacity first. Per model:
KV_pool = (GPU_memory × utilization) − weights − activation buffers; concurrency = pool ÷ (context × bytes/token). This is the fleet’s real capacity, in the units that matter. - Attack memory before hardware. fp8 KV cache, GQA/MLA architecture, quantized weights — each multiplies concurrency per card. Character.AI’s whole inference program is this step.
- Partition the static tier. MIG for tenant isolation (bet on a stable mix), time-slicing/MPS for long-tail density. One partitioning decision at provision time; re-partitions drain cards.
- Put the elastic tier on DRA claims with prioritized lists. Prefer the premium SKU, fall back to the cheaper one; let the scheduler finish your bin-packing (1.36+, NVIDIA DRA driver, GPUCluster mode).
- Scale on queue depth with an explicit tolerance band. HPA/KEDA on
vllm:num_requests_waiting(or TGI queue depth), scale-out threshold inside your SLO headroom, long cooldown for weight-load reality. - Re-derive capacity numbers monthly. Rates move (AWS cut P5 44% in June 2025); traffic duty cycles drift; model mix changes. A capacity plan older than a quarter is a guess with a spreadsheet skin.
When NOT to pack: the honest verdicts
- Flagship interactive serving: don’t share. Every packing technique taxes tail latency — time-slicing adds context-switch overhead, MPS shares SMs, over-commit adds swap-in latency (Cloudflare’s published 156ms cold-path). Your p95 SLO is the product; exclusive devices for the flagship tier, full stop.
- Training jobs: never pack. MIG slices cap memory and SMs; time-slicing destroys throughput determinism. Training economics are dominated by NVLink/IB topology anyway — packing training onto inference cards optimizes the wrong 5%.
- Fleets under ~8 GPUs: skip the machinery. DRA migration, DeviceClass tuning, KEDA thresholds, and MIG re-partition planning are fixed costs. Below ~8 GPUs the savings are hundreds of dollars against days of platform engineering — buy the idle capacity instead.
- Flat 24/7 traffic: skip elasticity. The elastic tier’s economics depend entirely on draining to zero off-peak. If traffic has no off-peak (global B2B, internal platforms), a static packed fleet is cheaper and simpler — the $17k/month figure above assumes a duty cycle that drains.
- Regulated/tenant-isolated environments: MIG or nothing. If your compliance story requires memory isolation between tenants, time-slicing/MPS/over-commit are all off the table — shared address spaces are shared audit surfaces.
References & further reading
- Kubernetes v1.36: More Drivers, New Features, and the Next Era of DRA — official blog, prioritized list graduation to stable
- Kubernetes v1.34: DRA has graduated to GA — the GA milestone post
- Kubernetes docs: Dynamic Resource Allocation — ResourceClaim/DeviceClass/ResourceSlice concepts
- DRA Driver for NVIDIA GPUs — SIG documentation for the production driver (GPU, MIG, VFIO, ComputeDomains)
- NVIDIA GPU Operator: DRA Driver for GPUs — install, prerequisites, alpha-status caveats, GPUCluster vs ClusterPolicy
- How Cloudflare runs more AI models on fewer GPUs — the Omni deep-dive (over-commit mechanics, 156ms swap)
- GKE: Best practices for autoscaling LLM inference — queue-size/batch-size guidance and DCGM metric limits
- Character.AI: Optimizing AI inference, part deux — int8 attention, cross-layer KV sharing
- vLLM metrics design — the Prometheus surface your HPA will scrape
- AWS: up to 45% price reduction for EC2 NVIDIA GPU instances — the June 2025 P5 cut
- Lambda GPU pricing — the specialist-tier rate card
- Kueue issue #15550 — DRA sharing quota admission tracking