vLLM Review: The Inference Engine That Turns GPU Utilization Into Your Problem
Sources
- vLLM v0.30.0 release notes
- vLLM documentation
- vLLM security model and API-key limitations
- vLLM scheduler and engine arguments (source, v0.30.0)
- Data-parallel deployment (vLLM docs)
- Automatic prefix caching design (vLLM docs)
- Sleep mode (vLLM docs)
- Per-request metrics (vLLM docs)
- vLLM benchmark CLI
- GuideLLM benchmarking framework
- NVIDIA H100 specification
- vLLM on PyPI
Every platform team with an AI roadmap eventually reaches the same slide: "GPU inference is our biggest cost line — let's self-host with vLLM and stop paying API margins." The vLLM project (Apache-2.0, now shipping v0.30.0 as the current stable) is the default answer, and on pure engineering it deserves to be: paged KV-cache management, continuous batching, prefix caching with salted hashes, and a V1 architecture that has quietly deleted most of the V0 era's operational sins. We read the v0.30.0 engine source end to end for this review — scheduler, block pool, admission control, authentication middleware — and built the cost model your finance team will eventually demand. The verdict is two-sided: vLLM is the best open inference server available, and self-hosting it is a utilization business that most teams are not staffed to win. The gap between those two statements is the entire review.
Executive Scorecard
| Dimension | Score | Why |
|---|---|---|
| Reliability | 7/10 | Engine core is solid; failure modes shift to fleet-level: preemption recompute storms, admission-control 503s, silent NaN pass-through by default |
| DX | 8/10 | OpenAI-compatible API, one-flag serving, first-class Prometheus metrics; but 200+ engine flags and device-dependent defaults make "tuned" configs non-portable |
| Cost | 4/10 (self-host) | $2,512/month idle per H100 at RunPod community rates; break-even vs cached API pricing demands ~24–48% sustained fleet utilization |
| Security | 5/10 | Built-in API-key auth guards only /v1, /v2, /inference, /cohere prefixes — /metrics, /health, and dev endpoints stay unauthenticated; prefix cache is a side channel unless you salt |
One scope note before the teardown: we did not run a GPU cluster for this review. Host-level GPU testing was not available in this environment, so every mechanical claim below is verified by reading the v0.30.0 source tree directly (file paths cited inline) and every cost number is computed from verified public rate cards — RunPod pricing, OpenAI pricing, Anthropic pricing, and DeepSeek pricing, all pulled live on 2026-09-28. Where a number depends on hardware behavior (tokens per second), we state the memory-bandwidth floor and label it as such. Nothing here is a benchmark claim.
Architecture Mechanics: What Actually Runs
The V1 engine is a pipeline of cooperating processes, not a monolith. Understanding the boundaries is prerequisite to operating it, because every interesting failure mode lives at one of them:
API server (FastAPI, async)
| HTTP/WebSocket, auth middleware, /metrics
| ZMQ over ipc:// or tcp://
v
EngineCore process(s) ---- per DP rank ----+
| Scheduler (event-driven) | DP rank 2..N
| KV cache manager / block pool | (mirror)
| |
v v
GPU worker(s) per rank (TP group)
- model runner: prefill/decode steps
- KV cache in HBM (paged blocks)
- sampler (seeded per-request RNG)
request lifecycle:
POST /v1/chat/completions
-> admission check (queue caps, 503 on overflow)
-> tokenization, block hashing (xxhash/sha256)
-> waiting queue -> schedule() per engine step
-> prefill chunks (chunked prefill) -> decode
-> detokenizer process -> SSE stream outThree mechanics matter more than the rest, so they get the deep treatment: how memory becomes concurrency, how the scheduler behaves under pressure, and what admission control does before pressure even starts.
The KV cache budget is your real concurrency limit
vLLM's headline invention — paged attention — treats KV memory like an OS treats virtual memory: fixed-size blocks, allocated on demand, shared across requests when block hashes collide. The engine sets aside a KV budget at startup as a fraction of GPU memory: gpu_memory_utilization defaults to 0.92 (vllm/config/cache.py, line 103), meaning a single instance assumes it owns 92% of the card. Whatever survives after model weights and activation buffers is your KV pool, and it is the number that decides how many conversations one GPU can hold. The per-token price of KV depends entirely on model architecture:
Dense GQA model (Qwen3-32B):
8 KV heads x 128 head_dim x 64 layers x 2 (K+V) x 2 bytes (bf16)
= 262,144 bytes/token = 256 KiB/token
fp8 KV (--kv-cache-dtype fp8) = 128 KiB/token
MLA model (DeepSeek-V3.2 class):
latent c_kv 512 + rope 64 = 576 bytes/token/layer, 61 layers, fp8
= 35,136 bytes/token = 34.3 KiB/token
= ~13% of the dense model's bf16 KV footprintNow the arithmetic that vendor GPU marketing never shows, for an 80 GB H100 at 0.92 utilization (79.0 GB usable; we count HBM in GiB and weights in decimal GB). Load Qwen3-32B with bf16 weights (65.6 GB) and your KV pool is 13.4 GB — 51,221 tokens total. A 10k-token conversation (8k context + 2k output) consumes 10,240 tokens of that pool. Concurrency: five conversations. The model advertises a 128k context window; on this configuration zero 132k-token conversations fit, and exactly one 30k conversation does. Quantize the weights to fp8 (32.8 GB) and the pool grows to 46.2 GB — 176k tokens, seventeen 10k conversations, exactly one 132k conversation. Add --kv-cache-dtype fp8 and you double it again (34 ten-k conversations, two 132k). Same GPU, same model, a 6.8x swing in concurrency purely from memory-format decisions:
| Configuration (Qwen3-32B on 1x H100 80GB) | KV pool | KV tokens | 10k-ctx convs | 132k-ctx convs |
|---|---|---|---|---|
| bf16 weights (65.6 GB), bf16 KV | 13.4 GB | 51,221 | 5 | 0 |
| fp8 weights (32.8 GB), bf16 KV | 46.2 GB | 176,344 | 17 | 1 |
| fp8 weights, fp8 KV | 46.2 GB | 352,687 | 34 | 2 |
This is the review's first hard finding: the "128k context" line in a model card is a spec, not a deployment property. Whether you can serve a long-context workload is decided by arithmetic like the above, not by the model. Platform teams that skip this step discover it empirically as OOM-loop restarts at 3 a.m.
Scheduler under pressure: preemption is recompute, not swap
When KV blocks run out mid-flight, the V1 scheduler preempts. The victim selection is FCFS inversion — it pops self.running[-1], the most recently admitted request, or the highest-priority request under priority scheduling (vllm/v1/core/sched/scheduler.py, lines ~745–790). Then comes the part that turns a scheduling hiccup into a capacity incident: _preempt_request frees all of the victim's blocks and sets request.num_computed_tokens = 0. There is no swap-out to host memory in V1 — the preemption mode flag from the V0 era is gone, and the only swap machinery left in the tree serves HiSparse, the sparse-MLA host-offload tier that shipped in 0.30.0. When the request resumes, it re-prefills from token zero.
The mitigation is prefix caching, and it is genuinely clever: freed cached blocks go to the tail of the free-block queue (FIFO, LRU eviction) while non-cached blocks go to the front (LIFO, reused first) — see block_pool.free_blocks. So a preempted request's blocks are usually still resident, and its re-prefill is mostly a hash lookup instead of a full forward pass. "Usually." Under sustained pressure, those cached blocks are exactly what gets evicted, and the recompute becomes real: a 28k-token conversation preempted at 90% completion re-prefills 28,672 tokens while the user stares at a frozen stream. The vllm:num_preemptions counter is on /metrics — alert on it, because preemption is the load shed you didn't configure.
0.30.0 also adds a knob for the thrash pattern: --watermark (SchedulerConfig field, default 0.0 = disabled) reserves a fraction of KV blocks as headroom before admitting waiting requests, trading a little capacity for admission stability. It works, but the default means most fleets discover it only after the incident.
Admission control: 503 before OOM
Newer than the scheduler itself is the API-server admission gate (vllm/v1/engine/async_llm.py, check_admission): --max-num-queued-reqs caps unfinished requests, --max-num-queued-tokens caps prompt tokens still in prefill. Overflow raises QueueOverflowError, which surfaces as HTTP 503 — deliberately, so load balancers retry on another replica. The accounting is conservative by design: prefix-cache hits and chunked-prefill progress are not subtracted from the token counter because the engine only reports them after prefill completes. That overestimation protects TTFT SLOs and costs you a few percent of admitted work. Set these caps; the default state on a fleet behind a round-robin LB is "first replica eats the whole queue."
And note the data-parallel trap documented in the DP deployment guide: --max-num-seqs is per rank, but --max-num-queued-reqs is per server — a --data-parallel-size=4 --max-num-seqs=256 --max-num-queued-reqs=256 deployment rejects requests once 256 are in flight total, even though the ranks could jointly run 1,024. The guide's own advice: size the queue cap to roughly data-parallel-size × max-num-seqs plus your queue depth.
Defaults are device-tier-dependent — your config is not portable
A quiet operational hazard: engine defaults are chosen per GPU tier at launch (vllm/engine/arg_utils.py, get_batch_defaults). H100/H200 (≥70 GB, non-A100) servers get max_num_batched_tokens=8192 and max_num_seqs=1024; B200-class (≥160 GB) gets 16384/1024; everything smaller — including A100s — drops to max_num_batched_tokens=2048 and max_num_seqs=256. The same launch command on two GPU generations is two different servers. Any benchmark or capacity plan produced on H100s silently overstates what an A100 fleet will do by 4x on batched tokens. Pin these flags in your deployment manifests and treat any doc that omits them as fiction.
Hands-On Breakdown: The Configuration Decisions That Matter
A production vLLM deployment that reflects the findings above — long-context serving on constrained memory — looks like this:
vllm serve Qwen/Qwen3-32B \
--quantization fp8 \
--kv-cache-dtype fp8 \
--max-model-len 131072 \
--max-num-seqs 64 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.92 \
--enable-prefix-caching \
--prefix-caching-hash-algo sha256_cbor \
--max-num-queued-reqs 128 \
--max-num-queued-tokens 262144 \
--watermark 0.05 \
--api-key "$VLLM_API_KEY" \
--served-model-name qwen3-32b-prodEvery flag traces to a finding: fp8 weights + fp8 KV is the only way a 132k conversation fits on one 80 GB card (34 ten-k conversations, two 132k); sha256_cbor over the default makes block hashes reproducible across Python and vLLM versions, which matters if you ever run disaggregated prefill or cross-engine cache sharing (the default pickle-based sha256 hashes are explicitly documented as not reproducible); queue caps + 503 admission protect TTFT before the scheduler starts preempting; the watermark trades 5% of capacity for admission stability. And the --api-key flag does less than you think — next section.
Multi-tenant isolation: cache_salt, the one flag nobody sets
Prefix caching is keyed by content hash. Without a salt, any request whose prefix hashes the same as another tenant's prefix will share blocks — an information side channel (a tenant can probe whether another tenant's prompt prefix exists by timing cache hits). vLLM's answer is the cache_salt request field (validated: non-empty string, ≤128 chars, no @/\ and NUL characters — vllm/entrypoints/generate/base/protocol.py): a per-tenant random value mixed into the first block's hash, which cascades down the whole prefix-hash chain. The field description in the OpenAI protocol layer recommends ~43 chars of base64 (256 bits). The engineering is sound. The deployment reality is that almost nobody sets it, because it requires per-request plumbing through your gateway. If you serve more than one trust domain on a replica, this is mandatory, not optional — and it is on you, not vLLM:
{
"model": "qwen3-32b-prod",
"messages": [{"role": "user", "content": "..."}],
"cache_salt": "tenant-7f3a-base64-43-char-random-value"
}The metrics that actually predict incidents
The Prometheus surface is first-class — full names verified in vllm/v1/metrics/loggers.py. Four of them predict incidents before users notice:
# KV exhaustion -> preemption storms
avg_over_time(vllm:num_preemptions[5m]) > 0
# Queue pressure building (503s imminent)
vllm:num_requests_waiting > on() group_left 0.75 * scalar(vllm:num_requests_running)
# Prefix cache hit collapse (multi-tenant churn)
rate(vllm:prefix_cache_hits[15m])
/ rate(vllm:prefix_cache_queries[15m]) < 0.3
# TTFT SLO erosion
histogram_quantile(0.99, rate(vllm:time_to_first_token_seconds_bucket[5m])) > 2One metric deserves a warning label: vllm:corrupted_requests only increments when VLLM_COMPUTE_NANS_IN_LOGITS is enabled (default false — vllm/envs.py line 249), because counting NaNs per step costs a reduction. Leave it off and a numerically unhappy model serves confidently wrong tokens with every green dashboard you have. If you run quantized checkpoints (and per the memory math above, you will), enable it and accept the overhead.
Critical Failure Modes
Ordered by how often we've seen them sink self-hosting pilots:
| Failure mode | Mechanism (verified in source) | Mitigation |
|---|---|---|
| Concurrency collapse on long contexts | KV budget = (92% × HBM) − weights; bf16 dense models leave single-digit concurrency; 132k conversations simply don't fit at bf16 | fp8 weights + fp8 KV; or MLA-architecture models (34.3 KiB/token vs 256); or multi-GPU TP |
| Preemption recompute storms | Preemption frees all blocks and resets num_computed_tokens=0 (scheduler.py:1477–1496); under pressure cached blocks evict too, so resume = full re-prefill |
--watermark 0.05; admission caps sized to DP width; alert on vllm:num_preemptions |
| Auth bypass on non-/v1 paths | AuthMiddleware guards only /v1, /v2, /inference, /cohere prefixes (authenticate.py:9); /metrics, /health, and dev routers (/sleep, /wake_up when VLLM_SERVER_DEV_MODE=1) are unauthenticated even with --api-key set |
Network-level ACLs; never expose the server directly; read vLLM's own security limitations doc before your security team does |
| Prefix cache side channel | Unsalted content-hash keys make cross-tenant prefix existence probe-able via timing | Per-tenant cache_salt at the gateway; 256-bit random values |
| Device-dependent defaults | Batched-token/seq defaults vary by GPU memory tier at launch (arg_utils.py:2708+); A100 fleets silently run 2048/256 where H100s ran 8192/1024 | Pin --max-num-batched-tokens and --max-num-seqs explicitly in manifests |
| Silent NaN corruption | vllm:corrupted_requests gated behind VLLM_COMPUTE_NANS_IN_LOGITS (default off) |
Enable the env var on quantized fleets; alert on the counter |
The Cost Math Nobody Runs Before Buying GPUs
Self-hosted inference is a utilization business. API vendors price tokens; you price time — and time is consumed 24/7 whether a request arrives or not. Below, the model is memory-bandwidth-floor decode throughput (an upper bound; real fleets run below it), on an H100 SXM (3.35 TB/s) at $3.49/hr (RunPod community, verified 2026-09-28), serving fp8 Qwen3-32B with fp8 KV. Per-step decode traffic is weights + batch × KV-per-token × context, so aggregate throughput falls as context grows — long contexts pay twice: once in KV memory, once in decode bandwidth.
Interactive: 28k ctx, batch 5 (fp8 weights + fp8 KV)
step traffic = 32.8 GB + 5 x 128 KiB x 28,672 = 51.6 GB/step
floor = 3350 GB/s / 51.6 GB = 65 tok/s/seq (aggregate 325 tok/s)
2k-token turn: ~32s wall; 5 turns decode concurrently
Batch: 8k ctx, batch 60
step traffic = 32.8 GB + 60 x 128 KiB x 8,192 = 97.2 GB/step
floor = 34.5 tok/s/seq (aggregate 2,067 tok/s)Turn the floors into money, against API prices pulled live on 2026-09-28 (gpt-6-sol and Claude Sonnet 5 at $2/$10 per 1M in/out with $0.20 cached input; DeepSeek V4.1-Flash at $0.30 miss / $0.006 hit / $1.20 out, peak rates):
| Scenario | Self-host (100% util) | Self-host (30% util) | gpt-6-sol / Sonnet 5, cached | DeepSeek Flash, cached |
|---|---|---|---|---|
| Agentic turn (28k ctx, 2k out) | $0.0061 | $0.0204 | $0.0262 | $0.0026 |
| Batch job (1,000 × 8k in / 2k out) | $1.49 | $4.97 | $36.86 | $4.92 |
Read the agentic row carefully, because it is the one that surprises teams: at 100% utilization self-hosting beats gpt-6-sol/Sonnet by ~4x, but the break-even point is ~23% sustained fleet utilization — and if your real throughput is half the memory floor (attention compute, sampling, kernel overheads), break-even rises to ~47%. Below that line, the API is cheaper than your own idle GPU time. The batch row shows the other edge: for high-batch offline work the API premium is 25x, which is why batch/synthetic-data teams almost always win by self-hosting. And DeepSeek's Flash pricing is the doomsday scenario for self-hosters of compact models — at $0.006/1M cached input it is below the raw GPU-time cost floor of hardware you own outright; no amount of utilization beats a vendor selling below your marginal cost.
Then there is the floor nobody escapes: an idle H100 at RunPod community rates is $2,512.80/month — a 4-GPU fleet is $10,051 — consumed whether or not a single request arrives. This is why the utilization number, not the per-token price, is the entire self-hosting decision. Add the operational stack you must now own (quantization validation, model rollout automation, capacity planning, on-call for GPU failures — the things an API absorbs) and the true cost of "cheap" self-hosted tokens climbs further.
Verdict
vLLM is the right engine and often the wrong decision. The engineering is genuinely excellent: paged KV management that survives hostile workloads, admission control that fails fast with 503s instead of OOM loops, prefix caching whose hash chain is well-designed and saltable, and a metrics surface that exposes everything an SRE actually needs. If you have made the decision to run inference infrastructure, vLLM is the correct choice — better-architected than the alternatives we've evaluated, and 0.30.0's sleep mode (POST /sleep?level=1 offloads weights to CPU and drops KV; level 2 discards everything for weight swaps) shows the project finally building for fleet economics, not just single-node throughput.
But the decision to run inference infrastructure is a utilization-business commitment, and that is a different commitment than the one most teams think they're making. The KV arithmetic in this review is one afternoon of work that prevents a quarter of pain: if your concurrency math produces single digits on your target hardware, no amount of vLLM tuning fixes it — the model or the memory format or the GPU count has to change.
Who should skip vLLM entirely
- Teams under ~25% projected GPU utilization — the math above says the API is cheaper than your idle time, including against premium frontier pricing.
- Anyone whose security posture requires a hardened HTTP server — vLLM's auth is a speed bump, not a gate; it guards four path prefixes and leaves /metrics and dev endpoints open even when configured. Put it behind a service mesh/gateway or don't put it on a network.
- Multi-tenant SaaS without gateway-level request plumbing — cache_salt isolation is real but manual; if you can't inject per-tenant salts today, you're one timing probe away from a cross-tenant disclosure conversation.
- Teams with zero GPU operations maturity — the failure modes here (preemption storms, device-dependent defaults, quantization NaNs) are all survivable with the observability vLLM ships and the runbooks this review sketches; without them you've bought a $2,500/month electric heater that occasionally hallucinates.
- Anyone chasing DeepSeek-Flash-class commodity pricing — $0.006/1M cached input is below the hardware cost floor; you cannot out-engineer a vendor selling below your marginal cost.
For everyone else — batch-heavy pipelines, latency-critical on-prem workloads, privacy-regulated inference, and teams with existing GPU fleets — vLLM 0.30.0 is the best answer on the market. Size your KV cache on paper before you size your fleet on a credit card.
References & Further Reading
- vLLM project repository — Apache-2.0; all source citations in this review are from the v0.30.0 tag
- vLLM security documentation — API-key authentication limitations
- Data-parallel deployment guide — per-rank vs per-server admission semantics
- Prefix caching design doc — hash-chain construction and cache_salt
- Per-request metrics and the vLLM benchmark CLI — including the cache-reuse warning that makes naive benchmark runs lie
- GuideLLM — the project's own recommended load-testing framework for production vLLM servers
- NVIDIA H100 datasheet — 80 GB HBM2e at 3.35 TB/s, the constant behind every throughput figure here
- Rate cards pulled live 2026-09-28: RunPod GPU pricing, OpenAI API pricing, Anthropic API pricing, DeepSeek API pricing — all subject to change; re-verify before modeling