GPU Capacity Planning for Inference Fleets: Bin-Packing H100s Without the 2 AM Page

Sources

GPU capacity planning for LLM inference is broken in most orgs, and the failure looks identical everywhere: a fleet of H100s provisioned for peak traffic sits at 5–15% utilization while a queue somewhere quietly breaches its SLO. The instinctive fix — buy more GPUs or shard harder — treats the symptom. The real problem is that capacity planning for inference is a memory problem disguised as a compute problem, and almost every tool in your stack treats it as the opposite.

This guide is the planning playbook that falls out of what production fleets actually do in 2026: Cloudflare packing 13 models onto a single GPU at ~400% memory over-commit, Character.AI redesigning attention before buying more hardware, and Kubernetes graduating DRA's prioritized device lists to stable in 1.36 so a claim can say “give me an H100, fall back to an A100 if none are free.” The order of operations matters: architecture first, partitioning second, elastic fleet last. Doing it in reverse is how you end up with a $200k/month fleet serving a product that never needed it.

TL;DR

The problem: GPUs are scheduled as integers, capacity is a memory curve

The Kubernetes device-plugin model that most fleets still run treats a GPU as an opaque countable integer: a pod either gets nvidia.com/gpu: 1 or nothing. That model has no way to express sharing, per-workload configuration, capability constraints, or topology requirements. Meanwhile the workload’s actual constraint is a memory curve: how many tokens the KV cache can hold for the model’s context length, which sets the concurrency ceiling of the card.

Run the numbers for a single H100 80GB serving a 32B dense GQA model in bf16. We computed the per-token KV cost and pool size for our vLLM review with Decimal-verified math; the shape of the result is what matters here:

Per-token KV bytes (dense GQA, 32B-class, 64 layers):
  8 KV heads x 128 head_dim x 64 layers x 2 (K+V) x 2 bytes
  = 262,144 bytes/token = 256 KiB/token

Usable H100 80GB memory pool at utilization 0.92:
  80 GiB x 0.92 ~= 73.7 GiB

  - model weights (32B x 2 bytes)      ~64.0 GiB   * does not fit; use fp8/AWQ
  - 8B-class model instead (16 GiB)     ~16.0 GiB
  ---------------------------------------
  KV pool for 8B model                 ~57.7 GiB

Concurrent conversations (8K context, 8K gen):
  57.7 GiB / (16,384 tokens x 256 KiB) ~= 14 concurrent streams

Three things fall out of that table, and they are the whole ballgame for capacity planning:

  1. Concurrency is KV_pool / (context × bytes_per_token), full stop. An 8B model on one H100 holds roughly 14 concurrent 16K-token conversations in bf16 with default preallocation. That is the real capacity of a $4–7/hour card, and it has nothing to do with FLOPs.
  2. Quantization is the first capacity lever, not the last. Cutting KV bytes (fp8 KV cache, GQA head pruning, or MLA architectures) multiplies concurrency at zero hardware cost. Character.AI’s inference stack (int8 attention with cross-layer KV sharing and hybrid attention horizons) exists precisely because “the KV cache is the bottleneck” — they redesigned attention before scaling the fleet.
  3. GPU utilization is the wrong signal at every level. GKE’s own autoscaling guidance says DCGM_FI_DEV_GPU_UTIL “does not measure how much work is being done while the GPU is active,” making it hard to map to latency; DCGM_FI_DEV_FB_USED won’t scale down preallocating servers like vLLM/TGI. More on the right signal below.

The planning unit that matters is tokens-resident: how many concurrent request-streams the fleet can hold at your p95 context length. Once you internalize that, every packing decision in this guide is arithmetic.

The 2026 packing toolkit: what changed and what it costs

For five years the answer to “can two workloads share a GPU?” was “time-slicing via the device plugin, and may God have mercy on your tail latency.” That changed. The current toolkit, in the order you should reach for it:

1. DRA: GPUs as claims, not integers

Dynamic Resource Allocation graduated to GA in Kubernetes 1.34 (September 2025), and the DynamicResourceAllocation feature gate is part of the default enabled set since 1.34/1.35 — on 1.32/1.33 you must enable it manually, per the NVIDIA DRA driver docs. The model is closer to PersistentVolumes than to device plugins: a workload declares what it needs in a ResourceClaim (possibly via a ResourceClaimTemplate), the driver publishes available devices as ResourceSlices, and a DeviceClass defines how a class of devices may be allocated — including sharing configuration that a device plugin could never express.

The graduation that matters for capacity planning landed in 1.36: prioritized lists went stable. From the official 1.36 DRA blog: instead of hardcoding a request for a specific device model, a claim can specify an ordered list of preferences — “give me an H100, but if none are available, fall back to an A100.” The scheduler evaluates preferences in order, which the release notes describe as drastically improving scheduling flexibility and cluster utilization. For a capacity planner, prioritized lists convert device heterogeneity from an ops headache into a bin-packing primitive: your premium tier claims H100s first and degrades gracefully; your batch tier claims A100s and only lands on H100s when they’d otherwise idle.

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: inference-gpu-preferred
spec:
  spec:
    resourceClassName: gpu.nvidia.com
    devices:
      requests:
      - name: gpu
        exactly:
          # first preference: 80GB Hopper with 700+ GB/s memory bandwidth
          - deviceClassName: gpu.nvidia.com
            selectors:
            - gpu.nvidia.com/model: H100
              gpu.nvidia.com/memory: "80"
              gpu.nvidia.com/brand: NVIDIA-H100-80GB
          # fallback: any 80GB card (A100 80GB acceptable)
          - deviceClassName: gpu.nvidia.com
            selectors:
            - gpu.nvidia.com/memory: "80"

Two operational caveats before you lean on this. First, driver maturity: the NVIDIA GPU Operator DRA driver lists some DRA features with alpha support status, requires Kubernetes v1.34.2+, and a cluster can run either the DRA-mode GPUCluster resource or the device-plugin ClusterPolicy — but not both. That is a real migration step, not a flag flip. Second, quota correctness: Kueue’s GPU admission control is still catching up with sharing — Kueue issue #15550 tracks quota-correct admission for DRA DeviceClasses with time-slicing/MPS sharing config. Until that lands, a Kueue quota counts claims, not slice-hours, and your chargeback math will disagree with your scheduler’s.

2. Time-slicing and MPS: the utilitarian packing tools

Below full-GPU claims, the NVIDIA DRA driver lets a single GPU allocation be shared across containers with time-slicing or Multi-Process Service (MPS), and lets one GPU be shared across independent claims via DRA “consumable capacity.” Time-slicing round-robins the GPU between contexts with no memory isolation; MPS gives concurrent kernel execution with a shared address space. Both raise packing density on underutilized cards; both blur your latency floor. The planning rule: time-slice your long-tail models (internal tools, batch enrichment, low-QPS endpoints) and keep flagship interactive serving on exclusive devices.

3. MIG: hardware partitioning when isolation is non-negotiable

MIG physically partitions an H100 into up to seven isolated instances with independent memory bandwidth and fault domains — a 1g.10gb slice is 10GB of memory and a guaranteed fraction of SMs. If one tenant’s runaway request can’t be allowed to OOM the neighbor’s inference pod, MIG is the only sharing mode that gives you that guarantee in hardware. The DRA driver manages MIG slices as first-class claimable devices, so a MIG slice can carry its own DeviceClass with per-class sharing config. The cost is flexibility: MIG partitioning is relatively static (a re-partition drains the card), so most fleets partition once at fleet-provisioning time based on the tenant mix they actually have.

H100 80GB, MIG profile 1g.10gb (1 GPU slice, 10GB GDDR):
  max 7x 1g.10gb slices per card (70GB usable; ~10GB reserved)

Tenant mix that motivated the partitioning:
  - 4x internal chatbot endpoints   (8B models, 4x 10GB slices)
  - 2x embedding/rerank services    (2x 10GB slices)
  - 1x canary/golden-path slot      (1x 10GB slice, kept free)

Fleet math:
  1 H100 -> 7 inference slots
  14 H100 fleet -> 98 slots -> serves ~40 internal models
  without MIG: same models need 40 GPUs (1 model/GPU)

That example is the entire MIG pitch: the same 14 cards serve 40 models instead of 14 at the cost of a one-time partitioning decision and a hard per-slice memory ceiling. The failure mode is equally simple — if a tenant’s model grows past 10GB, its slice OOMs at the partition boundary and no scheduler trick recovers it. MIG is a bet on a stable tenant mix.

4. Over-commit: Cloudflare’s Omni, and the latency you pay

The most aggressive packing data point in public is Cloudflare’s. Their Omni platform for Workers AI packs many small, low-volume models onto edge GPUs that would otherwise idle. The published mechanics: per-model process isolation with per-model Python virtual environments managed by uv, a CUDA stub library that intercepts cudaMalloc/cuMalloc calls and forces allocations into unified memory, and an overridden cudaMemGetInfo so each model only sees a subset of GPU memory. The result they publish: 13 models on a single GPU at ~400% memory over-commit, saving up to 4 GPUs per node in their deployment.

Over-commit works because inactive models migrate to CPU memory while active models stay hot on the GPU. The bill is a cold-start swap when a swapped-out model is called: moving a model with ~5GiB of weights and caches back onto the GPU takes ~156ms at PCIe 4.0’s ~32 GB/s. For Workers AI’s workload — thousands of small, low-QPS models where a sub-200ms cold-path is fine — that trade is obviously correct. For your flagship p50-sensitive endpoint, it is obviously wrong. Over-commit is a long-tail tool. Note also what Omni is not: it is not Kubernetes DRA, it is Cloudflare’s bespoke control plane — a reminder that the biggest packing wins still live below the orchestrator, in the runtime layer.

The fleet pattern: static base + elastic top, scheduled by queue depth

Packing decides how many models fit on a card. Fleet sizing decides how many cards you own. The pattern that production fleets converge on is a two-tier fleet:

                    +----------------------------+
  static base        |  N GPUs, always on         |   covers p50 demand at
  (MIG-partitioned,  |  MIG slices per tenant     |   24/7 duty cycle
   time-sliced)      +----------------------------+
                              ^
                              | overflow / peak
                              v
                    +----------------------------+
  elastic top        |  0..M GPUs, scale on       |   covers p95-p99 peaks,
  (exclusive GPUs,    |  queue depth (HPA +        |   drains to zero off-peak
   DRA claims)        |  KEDA on server metrics)   |

The static tier is where packing (MIG, time-slicing, prioritized claims) pays: it runs at high utilization because it is sized for the floor, not the peak. The elastic tier is where scheduling pays: it exists only to absorb peaks, and it should be exclusive-GPU (no packing) because peak traffic is exactly when you want zero contention. What triggers the elastic tier is the difference between a fleet that saves money and one that just adds a second idle fleet.

The right signal: queue size, not GPU utilization

Google’s GKE autoscaling best practices for LLM inference is the most direct public guidance on this, and it ranks the signals:

SignalWhat it measuresGoogle’s guidance
Queue size (server metric)Requests awaiting processing in the server queueUse to maximize throughput and cost within a target latency threshold; resilient to traffic fluctuation
Batch size (server metric)Requests currently undergoing inferenceUse to reach lower latency thresholds than queue size can
GPU utilization (DCGM_FI_DEV_GPU_UTIL)Duty cycle of GPU activityDoes not measure work done while active; hard to map to latency
GPU memory (DCGM_FI_DEV_FB_USED)Frame-buffer used at a point in timeOnly scales UP for preallocating servers (vLLM, TGI); never scales down
CPU/memoryGeneric utilizationNot recommended as sole indicators for GPU inference

The queue-size formula from the same doc: pick your target latency threshold, find the maximum throughput that holds it, then set the HPA target queue size to the level your fleet sustains at that throughput. And be mindful of the HPA tolerance, a default 0.1 no-action band around the target that dampens oscillation — set your target with that band in mind or your fleet will sawtooth. vLLM and TGI expose the queue and batch metrics natively; vLLM’s Prometheus surface is documented in its metrics design doc and we cover the operational details in our vLLM review.

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: inference-elastic-top
spec:
  scaleTargetRef:
    name: vllm-elastic
  minReplicaCount: 0    # elastic tier drains to zero off-peak
  maxReplicaCount: 8
  cooldownPeriod: 600    # 10 min: inference nodes drain slowly (weights load)
  triggers:
  - type: prometheus
    metadata:
      serverAddress: http://prometheus.monitoring:9090
      metricName: vllm:num_requests_waiting
      # threshold: requests waiting per replica before scale-out
      threshold: "4"
      query: avg(vllm:num_requests_waiting{deployment="vllm-elastic"})

Two production notes on that manifest. First, minReplicaCount: 0 plus cooldownPeriod: 600 encodes the real cost of scaling inference: a replica that scales up must load model weights (tens of seconds to minutes for large models), so you scale early on queue depth, not late. Second, the KEDA threshold is per-trigger; divide your fleet-wide queue target by replica count and re-derive it whenever the base tier’s packing changes.

The cost math: what packing and elasticity actually save (2026-09-28 rates)

Every number in this section is Decimal-verified against published rates on 2026-09-28 and rounded to cents. The input rates:

Provider / SKUConfigOn-demand ratePer-GPU
AWS p5.48xlarge8× H100 (SXM5, NVLink)$55.04/hr$6.88/hr
GCP a3-highgpu-8g8× H100$87.8325/hr$10.98/hr
Lambda 1× H100 SXM 80GB1 GPU, 208 vCPU$3.99/hr$3.99/hr

Context for the AWS number: AWS cut P5 on-demand prices 44% in June 2025 (from the May 31, 2025 baseline, per the announcement blog). Context for the spread: hyperscaler H100s cost 1.7–2.8× a specialist GPU cloud at on-demand rates, and that ratio is the entire economic argument for the two-tier fleet pattern — the elastic tier’s cost depends on who you can drain to zero fastest.

What 40% idle actually costs

64x H100 fleet, 40% of GPU-hours idle:

  AWS P5 per-GPU ($6.88/hr):    $4,227.07/day     $126,812.16/month
  GCP A3 per-GPU ($10.98/hr):   $6,745.54/day     $202,366.08/month
  Lambda per-GPU ($3.99/hr):     $2,451.46/day      $73,543.68/month

Same fleet, 20% idle (packing + elastic tier working):
  Lambda per-GPU:               $1,225.73/day      $36,771.84/month

Read that as: the difference between a packed-and-elastic fleet and a static one is roughly $36,772/month at Lambda rates for a 64-GPU fleet — and roughly $100k/month if the same fleet lives on GCP A3 on-demand. Packing is not a tuning exercise; at fleet scale it is a headcount-sized line item.

Static vs. elastic, same SLO

Static:   16 GPUs x $3.99 x 24h          = $1,532.16/day
Elastic:   10 avg GPUs x $3.99 x 24h       =   $957.60/day
                                          ----------
                                          $  574.56/day saved
                                          $17,236.80/month saved

The elastic fleet in that example averages 10 GPUs against a peak of 16 — the average comes from your traffic’s duty cycle, and the entire saving depends on the elastic tier draining to zero off-peak. If your traffic is flat around the clock (internal developer tools often are), elasticity saves nothing and the static packed fleet is the honest answer.

The interconnect tax on multi-node serving

When a model spans nodes (weights > single-card memory, or TP/EP sharding), interconnect bandwidth becomes a capacity-planning input, not a footnote. Time to move one node’s shard of a 70GB bf16 weight set (Decimal-verified):

PCIe 4.0 x16  ~32 GB/s    70 GB / 32 GB/s   = 2.188 s
NVLink (H100) ~900 GB/s   70 GB / 900 GB/s  = 0.078 s
400G IB       ~50 GB/s    70 GB / 50 GB/s   = 1.400 s

That 28× gap between PCIe and NVLink is why p5/a3 instances bundle NVSwitch and why “just shard across cheap single-GPU nodes” is usually false economy for models above ~40GB of weights. It is also why Cloudflare’s 156ms cold-swap number is tolerable at their scale (small models, PCIe-class edge hardware, latency-insensitive long-tail) and would be catastrophic on a flagship interactive endpoint.

Real deployments: who runs what

DeploymentPattern in productionPublished evidence
Cloudflare Workers AI (Omni)Per-model process isolation + CUDA-stub memory over-commit; 13 models/GPU at ~400% over-commit; cold-swap ~156ms for a 5GiB model over PCIe 4.0Cloudflare technical deep-dive
Character.AIArchitecture-first: int8 attention, cross-layer KV sharing, hybrid attention horizons — attack the KV bottleneck before scaling fleet sizeCharacter.AI part deux
Google (GKE guidance)Queue-size/batch-size HPA over GPU-utilization autoscaling; tolerance 0.1; DCGM metrics only for scale-upGKE autoscaling best practices
Kubernetes + NVIDIA DRAResourceClaim-first allocation with prioritized fallback (stable in 1.36); MIG/time-slicing/MPS as DeviceClass-configured sharingk8s 1.36 DRA blog, NVIDIA DRA driver docs

A pattern across all of them: nobody scaled their way out of a KV bottleneck. Cloudflare and Character.AI both attacked memory first, then packed, then bought hardware. The fleets that struggle are the ones that buy first.

The playbook, in order

  1. Compute your tokens-resident capacity first. Per model: KV_pool = (GPU_memory × utilization) − weights − activation buffers; concurrency = pool ÷ (context × bytes/token). This is the fleet’s real capacity, in the units that matter.
  2. Attack memory before hardware. fp8 KV cache, GQA/MLA architecture, quantized weights — each multiplies concurrency per card. Character.AI’s whole inference program is this step.
  3. Partition the static tier. MIG for tenant isolation (bet on a stable mix), time-slicing/MPS for long-tail density. One partitioning decision at provision time; re-partitions drain cards.
  4. Put the elastic tier on DRA claims with prioritized lists. Prefer the premium SKU, fall back to the cheaper one; let the scheduler finish your bin-packing (1.36+, NVIDIA DRA driver, GPUCluster mode).
  5. Scale on queue depth with an explicit tolerance band. HPA/KEDA on vllm:num_requests_waiting (or TGI queue depth), scale-out threshold inside your SLO headroom, long cooldown for weight-load reality.
  6. Re-derive capacity numbers monthly. Rates move (AWS cut P5 44% in June 2025); traffic duty cycles drift; model mix changes. A capacity plan older than a quarter is a guess with a spreadsheet skin.

When NOT to pack: the honest verdicts

References & further reading