vLLM Review: The Inference Engine That Turns GPU Utilization Into Your Problem

Sources

Every platform team with an AI roadmap eventually reaches the same slide: "GPU inference is our biggest cost line — let's self-host with vLLM and stop paying API margins." The vLLM project (Apache-2.0, now shipping v0.30.0 as the current stable) is the default answer, and on pure engineering it deserves to be: paged KV-cache management, continuous batching, prefix caching with salted hashes, and a V1 architecture that has quietly deleted most of the V0 era's operational sins. We read the v0.30.0 engine source end to end for this review — scheduler, block pool, admission control, authentication middleware — and built the cost model your finance team will eventually demand. The verdict is two-sided: vLLM is the best open inference server available, and self-hosting it is a utilization business that most teams are not staffed to win. The gap between those two statements is the entire review.

Executive Scorecard

DimensionScoreWhy
Reliability7/10Engine core is solid; failure modes shift to fleet-level: preemption recompute storms, admission-control 503s, silent NaN pass-through by default
DX8/10OpenAI-compatible API, one-flag serving, first-class Prometheus metrics; but 200+ engine flags and device-dependent defaults make "tuned" configs non-portable
Cost4/10 (self-host)$2,512/month idle per H100 at RunPod community rates; break-even vs cached API pricing demands ~24–48% sustained fleet utilization
Security5/10Built-in API-key auth guards only /v1, /v2, /inference, /cohere prefixes — /metrics, /health, and dev endpoints stay unauthenticated; prefix cache is a side channel unless you salt

One scope note before the teardown: we did not run a GPU cluster for this review. Host-level GPU testing was not available in this environment, so every mechanical claim below is verified by reading the v0.30.0 source tree directly (file paths cited inline) and every cost number is computed from verified public rate cards — RunPod pricing, OpenAI pricing, Anthropic pricing, and DeepSeek pricing, all pulled live on 2026-09-28. Where a number depends on hardware behavior (tokens per second), we state the memory-bandwidth floor and label it as such. Nothing here is a benchmark claim.

Architecture Mechanics: What Actually Runs

The V1 engine is a pipeline of cooperating processes, not a monolith. Understanding the boundaries is prerequisite to operating it, because every interesting failure mode lives at one of them:

   API server (FastAPI, async)
        |  HTTP/WebSocket, auth middleware, /metrics
        |  ZMQ over ipc:// or tcp://
        v
   EngineCore process(s)  ---- per DP rank ----+
        |   Scheduler (event-driven)          |  DP rank 2..N
        |   KV cache manager / block pool      |  (mirror)
        |                                      |
        v                                      v
   GPU worker(s) per rank (TP group)
        - model runner: prefill/decode steps
        - KV cache in HBM (paged blocks)
        - sampler (seeded per-request RNG)

  request lifecycle:
  POST /v1/chat/completions
    -> admission check (queue caps, 503 on overflow)
    -> tokenization, block hashing (xxhash/sha256)
    -> waiting queue -> schedule() per engine step
    -> prefill chunks (chunked prefill) -> decode
    -> detokenizer process -> SSE stream out

Three mechanics matter more than the rest, so they get the deep treatment: how memory becomes concurrency, how the scheduler behaves under pressure, and what admission control does before pressure even starts.

The KV cache budget is your real concurrency limit

vLLM's headline invention — paged attention — treats KV memory like an OS treats virtual memory: fixed-size blocks, allocated on demand, shared across requests when block hashes collide. The engine sets aside a KV budget at startup as a fraction of GPU memory: gpu_memory_utilization defaults to 0.92 (vllm/config/cache.py, line 103), meaning a single instance assumes it owns 92% of the card. Whatever survives after model weights and activation buffers is your KV pool, and it is the number that decides how many conversations one GPU can hold. The per-token price of KV depends entirely on model architecture:

Dense GQA model (Qwen3-32B):
  8 KV heads x 128 head_dim x 64 layers x 2 (K+V) x 2 bytes (bf16)
  = 262,144 bytes/token = 256 KiB/token
  fp8 KV (--kv-cache-dtype fp8) = 128 KiB/token

MLA model (DeepSeek-V3.2 class):
  latent c_kv 512 + rope 64 = 576 bytes/token/layer, 61 layers, fp8
  = 35,136 bytes/token = 34.3 KiB/token
  = ~13% of the dense model's bf16 KV footprint

Now the arithmetic that vendor GPU marketing never shows, for an 80 GB H100 at 0.92 utilization (79.0 GB usable; we count HBM in GiB and weights in decimal GB). Load Qwen3-32B with bf16 weights (65.6 GB) and your KV pool is 13.4 GB — 51,221 tokens total. A 10k-token conversation (8k context + 2k output) consumes 10,240 tokens of that pool. Concurrency: five conversations. The model advertises a 128k context window; on this configuration zero 132k-token conversations fit, and exactly one 30k conversation does. Quantize the weights to fp8 (32.8 GB) and the pool grows to 46.2 GB — 176k tokens, seventeen 10k conversations, exactly one 132k conversation. Add --kv-cache-dtype fp8 and you double it again (34 ten-k conversations, two 132k). Same GPU, same model, a 6.8x swing in concurrency purely from memory-format decisions:

Configuration (Qwen3-32B on 1x H100 80GB)KV poolKV tokens10k-ctx convs132k-ctx convs
bf16 weights (65.6 GB), bf16 KV13.4 GB51,22150
fp8 weights (32.8 GB), bf16 KV46.2 GB176,344171
fp8 weights, fp8 KV46.2 GB352,687342

This is the review's first hard finding: the "128k context" line in a model card is a spec, not a deployment property. Whether you can serve a long-context workload is decided by arithmetic like the above, not by the model. Platform teams that skip this step discover it empirically as OOM-loop restarts at 3 a.m.

Scheduler under pressure: preemption is recompute, not swap

When KV blocks run out mid-flight, the V1 scheduler preempts. The victim selection is FCFS inversion — it pops self.running[-1], the most recently admitted request, or the highest-priority request under priority scheduling (vllm/v1/core/sched/scheduler.py, lines ~745–790). Then comes the part that turns a scheduling hiccup into a capacity incident: _preempt_request frees all of the victim's blocks and sets request.num_computed_tokens = 0. There is no swap-out to host memory in V1 — the preemption mode flag from the V0 era is gone, and the only swap machinery left in the tree serves HiSparse, the sparse-MLA host-offload tier that shipped in 0.30.0. When the request resumes, it re-prefills from token zero.

The mitigation is prefix caching, and it is genuinely clever: freed cached blocks go to the tail of the free-block queue (FIFO, LRU eviction) while non-cached blocks go to the front (LIFO, reused first) — see block_pool.free_blocks. So a preempted request's blocks are usually still resident, and its re-prefill is mostly a hash lookup instead of a full forward pass. "Usually." Under sustained pressure, those cached blocks are exactly what gets evicted, and the recompute becomes real: a 28k-token conversation preempted at 90% completion re-prefills 28,672 tokens while the user stares at a frozen stream. The vllm:num_preemptions counter is on /metrics — alert on it, because preemption is the load shed you didn't configure.

0.30.0 also adds a knob for the thrash pattern: --watermark (SchedulerConfig field, default 0.0 = disabled) reserves a fraction of KV blocks as headroom before admitting waiting requests, trading a little capacity for admission stability. It works, but the default means most fleets discover it only after the incident.

Admission control: 503 before OOM

Newer than the scheduler itself is the API-server admission gate (vllm/v1/engine/async_llm.py, check_admission): --max-num-queued-reqs caps unfinished requests, --max-num-queued-tokens caps prompt tokens still in prefill. Overflow raises QueueOverflowError, which surfaces as HTTP 503 — deliberately, so load balancers retry on another replica. The accounting is conservative by design: prefix-cache hits and chunked-prefill progress are not subtracted from the token counter because the engine only reports them after prefill completes. That overestimation protects TTFT SLOs and costs you a few percent of admitted work. Set these caps; the default state on a fleet behind a round-robin LB is "first replica eats the whole queue."

And note the data-parallel trap documented in the DP deployment guide: --max-num-seqs is per rank, but --max-num-queued-reqs is per server — a --data-parallel-size=4 --max-num-seqs=256 --max-num-queued-reqs=256 deployment rejects requests once 256 are in flight total, even though the ranks could jointly run 1,024. The guide's own advice: size the queue cap to roughly data-parallel-size × max-num-seqs plus your queue depth.

Defaults are device-tier-dependent — your config is not portable

A quiet operational hazard: engine defaults are chosen per GPU tier at launch (vllm/engine/arg_utils.py, get_batch_defaults). H100/H200 (≥70 GB, non-A100) servers get max_num_batched_tokens=8192 and max_num_seqs=1024; B200-class (≥160 GB) gets 16384/1024; everything smaller — including A100s — drops to max_num_batched_tokens=2048 and max_num_seqs=256. The same launch command on two GPU generations is two different servers. Any benchmark or capacity plan produced on H100s silently overstates what an A100 fleet will do by 4x on batched tokens. Pin these flags in your deployment manifests and treat any doc that omits them as fiction.

Hands-On Breakdown: The Configuration Decisions That Matter

A production vLLM deployment that reflects the findings above — long-context serving on constrained memory — looks like this:

vllm serve Qwen/Qwen3-32B \
  --quantization fp8 \
  --kv-cache-dtype fp8 \
  --max-model-len 131072 \
  --max-num-seqs 64 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.92 \
  --enable-prefix-caching \
  --prefix-caching-hash-algo sha256_cbor \
  --max-num-queued-reqs 128 \
  --max-num-queued-tokens 262144 \
  --watermark 0.05 \
  --api-key "$VLLM_API_KEY" \
  --served-model-name qwen3-32b-prod

Every flag traces to a finding: fp8 weights + fp8 KV is the only way a 132k conversation fits on one 80 GB card (34 ten-k conversations, two 132k); sha256_cbor over the default makes block hashes reproducible across Python and vLLM versions, which matters if you ever run disaggregated prefill or cross-engine cache sharing (the default pickle-based sha256 hashes are explicitly documented as not reproducible); queue caps + 503 admission protect TTFT before the scheduler starts preempting; the watermark trades 5% of capacity for admission stability. And the --api-key flag does less than you think — next section.

Multi-tenant isolation: cache_salt, the one flag nobody sets

Prefix caching is keyed by content hash. Without a salt, any request whose prefix hashes the same as another tenant's prefix will share blocks — an information side channel (a tenant can probe whether another tenant's prompt prefix exists by timing cache hits). vLLM's answer is the cache_salt request field (validated: non-empty string, ≤128 chars, no @/\ and NUL characters — vllm/entrypoints/generate/base/protocol.py): a per-tenant random value mixed into the first block's hash, which cascades down the whole prefix-hash chain. The field description in the OpenAI protocol layer recommends ~43 chars of base64 (256 bits). The engineering is sound. The deployment reality is that almost nobody sets it, because it requires per-request plumbing through your gateway. If you serve more than one trust domain on a replica, this is mandatory, not optional — and it is on you, not vLLM:

{
  "model": "qwen3-32b-prod",
  "messages": [{"role": "user", "content": "..."}],
  "cache_salt": "tenant-7f3a-base64-43-char-random-value"
}

The metrics that actually predict incidents

The Prometheus surface is first-class — full names verified in vllm/v1/metrics/loggers.py. Four of them predict incidents before users notice:

# KV exhaustion -> preemption storms
avg_over_time(vllm:num_preemptions[5m]) > 0

# Queue pressure building (503s imminent)
vllm:num_requests_waiting > on() group_left 0.75 * scalar(vllm:num_requests_running)

# Prefix cache hit collapse (multi-tenant churn)
rate(vllm:prefix_cache_hits[15m])
  / rate(vllm:prefix_cache_queries[15m]) < 0.3

# TTFT SLO erosion
histogram_quantile(0.99, rate(vllm:time_to_first_token_seconds_bucket[5m])) > 2

One metric deserves a warning label: vllm:corrupted_requests only increments when VLLM_COMPUTE_NANS_IN_LOGITS is enabled (default false — vllm/envs.py line 249), because counting NaNs per step costs a reduction. Leave it off and a numerically unhappy model serves confidently wrong tokens with every green dashboard you have. If you run quantized checkpoints (and per the memory math above, you will), enable it and accept the overhead.

Critical Failure Modes

Ordered by how often we've seen them sink self-hosting pilots:

Failure modeMechanism (verified in source)Mitigation
Concurrency collapse on long contexts KV budget = (92% × HBM) − weights; bf16 dense models leave single-digit concurrency; 132k conversations simply don't fit at bf16 fp8 weights + fp8 KV; or MLA-architecture models (34.3 KiB/token vs 256); or multi-GPU TP
Preemption recompute storms Preemption frees all blocks and resets num_computed_tokens=0 (scheduler.py:1477–1496); under pressure cached blocks evict too, so resume = full re-prefill --watermark 0.05; admission caps sized to DP width; alert on vllm:num_preemptions
Auth bypass on non-/v1 paths AuthMiddleware guards only /v1, /v2, /inference, /cohere prefixes (authenticate.py:9); /metrics, /health, and dev routers (/sleep, /wake_up when VLLM_SERVER_DEV_MODE=1) are unauthenticated even with --api-key set Network-level ACLs; never expose the server directly; read vLLM's own security limitations doc before your security team does
Prefix cache side channel Unsalted content-hash keys make cross-tenant prefix existence probe-able via timing Per-tenant cache_salt at the gateway; 256-bit random values
Device-dependent defaults Batched-token/seq defaults vary by GPU memory tier at launch (arg_utils.py:2708+); A100 fleets silently run 2048/256 where H100s ran 8192/1024 Pin --max-num-batched-tokens and --max-num-seqs explicitly in manifests
Silent NaN corruption vllm:corrupted_requests gated behind VLLM_COMPUTE_NANS_IN_LOGITS (default off) Enable the env var on quantized fleets; alert on the counter

The Cost Math Nobody Runs Before Buying GPUs

Self-hosted inference is a utilization business. API vendors price tokens; you price time — and time is consumed 24/7 whether a request arrives or not. Below, the model is memory-bandwidth-floor decode throughput (an upper bound; real fleets run below it), on an H100 SXM (3.35 TB/s) at $3.49/hr (RunPod community, verified 2026-09-28), serving fp8 Qwen3-32B with fp8 KV. Per-step decode traffic is weights + batch × KV-per-token × context, so aggregate throughput falls as context grows — long contexts pay twice: once in KV memory, once in decode bandwidth.

Interactive: 28k ctx, batch 5 (fp8 weights + fp8 KV)
  step traffic = 32.8 GB + 5 x 128 KiB x 28,672 = 51.6 GB/step
  floor = 3350 GB/s / 51.6 GB = 65 tok/s/seq (aggregate 325 tok/s)
  2k-token turn: ~32s wall; 5 turns decode concurrently

Batch: 8k ctx, batch 60
  step traffic = 32.8 GB + 60 x 128 KiB x 8,192 = 97.2 GB/step
  floor = 34.5 tok/s/seq (aggregate 2,067 tok/s)

Turn the floors into money, against API prices pulled live on 2026-09-28 (gpt-6-sol and Claude Sonnet 5 at $2/$10 per 1M in/out with $0.20 cached input; DeepSeek V4.1-Flash at $0.30 miss / $0.006 hit / $1.20 out, peak rates):

ScenarioSelf-host (100% util)Self-host (30% util)gpt-6-sol / Sonnet 5, cachedDeepSeek Flash, cached
Agentic turn (28k ctx, 2k out)$0.0061$0.0204$0.0262$0.0026
Batch job (1,000 × 8k in / 2k out)$1.49$4.97$36.86$4.92

Read the agentic row carefully, because it is the one that surprises teams: at 100% utilization self-hosting beats gpt-6-sol/Sonnet by ~4x, but the break-even point is ~23% sustained fleet utilization — and if your real throughput is half the memory floor (attention compute, sampling, kernel overheads), break-even rises to ~47%. Below that line, the API is cheaper than your own idle GPU time. The batch row shows the other edge: for high-batch offline work the API premium is 25x, which is why batch/synthetic-data teams almost always win by self-hosting. And DeepSeek's Flash pricing is the doomsday scenario for self-hosters of compact models — at $0.006/1M cached input it is below the raw GPU-time cost floor of hardware you own outright; no amount of utilization beats a vendor selling below your marginal cost.

Then there is the floor nobody escapes: an idle H100 at RunPod community rates is $2,512.80/month — a 4-GPU fleet is $10,051 — consumed whether or not a single request arrives. This is why the utilization number, not the per-token price, is the entire self-hosting decision. Add the operational stack you must now own (quantization validation, model rollout automation, capacity planning, on-call for GPU failures — the things an API absorbs) and the true cost of "cheap" self-hosted tokens climbs further.

Verdict

vLLM is the right engine and often the wrong decision. The engineering is genuinely excellent: paged KV management that survives hostile workloads, admission control that fails fast with 503s instead of OOM loops, prefix caching whose hash chain is well-designed and saltable, and a metrics surface that exposes everything an SRE actually needs. If you have made the decision to run inference infrastructure, vLLM is the correct choice — better-architected than the alternatives we've evaluated, and 0.30.0's sleep mode (POST /sleep?level=1 offloads weights to CPU and drops KV; level 2 discards everything for weight swaps) shows the project finally building for fleet economics, not just single-node throughput.

But the decision to run inference infrastructure is a utilization-business commitment, and that is a different commitment than the one most teams think they're making. The KV arithmetic in this review is one afternoon of work that prevents a quarter of pain: if your concurrency math produces single digits on your target hardware, no amount of vLLM tuning fixes it — the model or the memory format or the GPU count has to change.

Who should skip vLLM entirely

For everyone else — batch-heavy pipelines, latency-critical on-prem workloads, privacy-regulated inference, and teams with existing GPU fleets — vLLM 0.30.0 is the best answer on the market. Size your KV cache on paper before you size your fleet on a credit card.

References & Further Reading