Model-Serving Topologies After Beam: When 501B Parameters Break Your Single-Cluster Instincts
Sources
- Reflection: Introducing Beam — 501B sparse MoE, 23B active (Oct 5, 2026)
- DeepSeek open-infra-index: V3/R1 Inference System Overview (Day 6)
- DeepSeek-V3 Technical Report (arXiv 2412.19437) — deployment sections
- SGLang docs: Expert Parallelism
- SGLang docs: PD Disaggregation
- vLLM docs: Parallelism and Scaling
- vLLM docs: Data Parallel Deployment
- NVIDIA Dynamo docs: Disaggregated Serving overview
- Slack Engineering: Migration to a Cellular Architecture
- AWS Well-Architected: Reducing the Scope of Impact with Cell-Based Architecture
- RunPod GPU pricing (pulled live 2026-10-06)
- llm-d: distributed inference on Kubernetes
- AIBrix: production LLM inference orchestration on K8s
Reflection shipped Beam on October 5, 2026: a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, trained on 23.8T tokens, with a reinforcement-learning run that burned 10,500 NVIDIA GB300 GPUs for four weeks to generate over 100 million rollouts. The weights are not public yet — Reflection promises the “weights, technical report, model card, and developer artifacts later this month.” That is the story. Not the benchmarks.
The benchmarks say Beam is competitive with GLM-5.2 while using 3–4× less inference compute, and it approaches Qwen 3.8-Max on coding and agentic tasks. Kimi K3 remains ahead on raw capability. Reflection’s own framing is honest here: the pitch is efficiency at inference time, not frontier supremacy. A 501B/23B sparse MoE at Western-open weights, arriving weeks after Kimi K3 and GLM-5.3, means the open-weight 500B class is now a routine purchase decision — and it breaks the single-cluster instincts most platform teams built around 70B-class models.
This guide is the serving-topology briefing for that world. The short version: a Beam-class model in FP8 “fits” on 8 H100s with about 11 GB left for KV cache. Fits is not serves. Everything below is about what actually serves.
The 501B/23B Trap: Fits-Is-Not-Serves
The arithmetic that catches teams is simple and merciless. Beam’s weights, by format:
BF16 (2 bytes/param): 501e9 * 2 = 1,002 GB across the fleet
FP8 (1 byte/param): 501e9 * 1 = 501 GB across the fleet
H100 80GB, BF16: ceil(1002/80) = 13 GPUs for weights alone
H100 80GB, FP8: ceil(501/80) = 7 GPUs for weights alone
H200 141GB, FP8: ceil(501/141) = 4 GPUs for weights alone
At 80% usable VRAM (weights + KV + activations + NCCL buffers):
H100 FP8: ceil(501/(80*0.8)) = 8 GPUs, spare = 8*64-501 = 11 GB fleet-wide
Eleven gigabytes. A Beam-class GQA model with roughly 60 layers, 8 KV heads, 128-dim heads at FP8 KV eats ~120 KiB per token. Eleven GB buys 0.09M tokens of KV — at 32k context that is three concurrent requests across the entire replica, before activations. The “it fits” cluster is a machine that can hold the model and almost nothing else.
This is not a Beam defect. It is the defining property of the sparse-MoE class since DeepSeek-V3 (671B total / 37B active): weights memory scales with total parameters; compute per token scales with active parameters. The two axes are decoupled, and every topology decision below is a consequence of that decoupling. You buy memory for 501B and FLOPs for 23B. The 23B active side is genuinely cheap — Beam’s forward pass is ~46 GFLOPs/token, dense-23B arithmetic — which is why Reflection can claim GLM-5.2-class reasoning at a fraction of the compute. The 501B side is the bill.
The fix for the KV squeeze is never “more GPUs in the same replica.” Adding GPUs to a tensor-parallel replica shrinks per-GPU KV slightly (each rank holds a shard), but NCCL all-reduce communication grows and per-GPU KV still trails far behind what a serving workload needs. The fix is more replicas — which is where topology starts to matter.
Topology 1: The Aggregated Single-Cluster Default (And Its Ceiling)
The default topology every team starts with: one Kubernetes cluster, one vLLM or SGLang deployment, TP=8 per node, scale replicas horizontally under an HPA or llm-d-style gateway. It works because at 70B-class, a whole replica fits in one node’s NVLink domain and a single pod is the unit of both capacity and failure.
Beam-class models strain both properties at once:
- Weight placement forces multi-node replicas. FP8 needs 7 GPUs minimum; BF16 needs 13. Either way you cross node boundaries, and cross-node tensor parallelism lives or dies by network fabric — vLLM’s guidance is explicit: TP within a node, PP across nodes, and pipeline parallelism preferred when GPUs lack NVLink. The moment your replica spans nodes, the pod is no longer your failure domain; the rank group is.
- Expert parallelism changes the traffic pattern. MoE layers do not all-reduce; they all-to-all. Every token, every layer, must reach the GPUs hosting its chosen experts. vLLM supports data-parallel attention with expert-parallel MoE; SGLang exposes it via
--moe-a2a-backend deepepand--epflags with the same constraint:ep_size = tp_sizefor the high-performance backends. - The single cluster becomes a shared-fate domain. One bad node in the rank group stalls the replica; one bad scheduler upgrade stalls the fleet.
The ceiling is not throughput. It is blast radius and utilization: a single-cluster fleet of expensive multi-node replicas is either sized for peak (idle money) or sized for average (queueing at peak). DeepSeek’s published fleet behavior shows the alternative.
Topology 2: Disaggregated Prefill/Decode — The DeepSeek Pattern
The most consequential production artifact for MoE serving is DeepSeek’s inference system overview (open-infra-index, Day 6). They run prefill and decode as separate deployments with different parallelism shapes, because the two phases have different bottlenecks:
Prefill unit: 4 nodes, 32 GPUs — routed-expert EP32, 9 routed experts
+ 1 shared expert per GPU, 32 redundant experts
Decode unit: 18 nodes, 144 GPUs — routed-expert EP144, 2 routed experts
+ 1 shared expert per GPU, 32 redundant experts
Why different: prefill is compute-bound (big batches, all-to-all hidden
behind GEMMs); decode is memory-bound (small per-expert
batches, all-to-all latency exposed on every token)
The decode unit is 4.5× the prefill unit in GPUs. That asymmetry is the whole point: prefill throughput scales with compute, decode throughput scales with KV and expert placement, and forcing them into one pool means one phase’s SLO is hostage to the other’s tail. Their numbers from a 24h production window (Feb 27–28, 2025): peak 278 nodes, average 226.75 nodes (8 H800/node), 608B input tokens (56.3% hitting the on-disk KV cache), 168B output tokens, ~73.7k input tok/s per node prefilling, ~14.8k output tok/s per node decoding, at an assumed $2/hr H800 lease = $87,072/day. Two load balancers — one for prefill (balances core-attention compute and dispatch send load), one for decode (balances KV usage and request counts) — plus an expert-parallel load balancer that detects high-load experts and re-replicates them every ~10 minutes. Nighttime, they shrink the inference fleet and hand nodes to research and training.
That daily flex is the pattern platform engineers should steal: inference capacity is a schedule, not a constant. Your 40% idle at 03:00 is someone else’s training capacity. DeepSeek’s own cost floor from those numbers is $0.518 per 1M output tokens at $2/hr — compare with your API bill before you build any of this.
The OSS replication path for this pattern: SGLang’s PD-disaggregation mode, or NVIDIA Dynamo’s DynamoGraphDeployment, whose disaggregated YAML is a three-line diff from aggregated:
services:
Frontend:
componentType: frontend
prefill:
componentType: worker
subComponentType: prefill # prompt processing only
decode:
componentType: worker
subComponentType: decode # token generation only
Dynamo’s docs are refreshingly blunt about when not to do this: “It is not automatically better. For small models, short prompts, low concurrency, or clusters without a fast KV-transfer fabric, an aggregated deployment is simpler and often faster.” The KV-transfer fabric is the hidden cost — prefill and decode pools only work if KV moves between them faster than it can be recomputed, which means RDMA or NVLink, which means network planning before the first pod lands.
Topology 3: Cell-Based Serving — Bounding the Blast Radius
The third topology is not about performance at all. It is about failure. A multi-node TP/EP replica is a shared-fate unit: lose one node and the replica’s KV cache is lost with it. Once your fleet is a handful of expensive replicas, the failure domain is your whole inference product.
The fix is the cell pattern, borrowed straight from general cloud architecture: Slack’s cellular migration made AZs drainable cells with an Envoy/xDS edge reweighting traffic; AWS’s Well-Architected cell guidance formalizes it: a thin router, N isolated cells, each cell handling a bounded slice of tenants. Applied to serving:
- Cell = one disaggregated deployment unit. Using DeepSeek’s shapes: 4-node prefill + 18-node decode = 22 nodes per cell. At RunPod’s 2026-10-06 rates ($2.69/hr H100 SXM Community), that cell costs $473/hr; Secure Cloud $3.49 → $614/hr.
- Router = your existing gateway. AIBrix and llm-d both implement prefix-aware, KV-aware routing with per-cell targeting; nothing exotic is required beyond stable cell IDs in the router config.
- Blast radius = one cell. Bad weights rollout? One cell takes the hit; you drain it and re-point tenants via the router while the others serve. Bad expert distribution after a traffic shift? It is contained inside one cell’s EPLB loop.
The cost is honesty about redundancy: N+1 cell redundancy (2 active + 1 warm standby) is a 50% hardware premium over N cells. For a 22-node cell fleet at RunPod Community rates, warm-standby redundancy adds $473/hr — about $346k/month. You buy that only when the product is enterprise-facing enough to need bounded blast radius, or when your cells are large enough that losing one is a page rather than an incident.
Picking a Topology: The Decision Table
| Topology | Replica shape | Best fit | Cost driver | When NOT to use |
|---|---|---|---|---|
| Aggregated single-cluster | TP=8 per node, HPA replicas | ≤70B dense or small MoE; POCs; teams without RDMA | GPU-hours × utilization | Beam-class MoE at FP8 floor (11GB KV headroom); multi-region SLOs |
| Disaggregated prefill/decode | 4-node prefill + 18-node decode units (DeepSeek shapes) | Reasoning/agentic traffic (long prompts, long generations); ≥95% decode SLO needs | KV-transfer fabric + node count | Short prompts, low concurrency, no RDMA — aggregated is simpler and often faster (Dynamo docs) |
| Cell-based multi-cell | N × (prefill+decode unit), thin router | Enterprise multi-tenant; bounded-blast-radius requirements | N+1 redundancy premium (~50%) | <3 cells of demand — the router + control plane tax exceeds the resilience win |
Read the table as cumulative, not exclusive: cells contain disaggregated units, which contain TP/EP replicas. The mistake is skipping levels — running cells without disaggregation wastes the cell boundary, running disaggregation without cells leaves the biggest failure domain untouched.
Cost Math That Survives a Finance Review
Order-of-magnitude economics for a Beam-class FP8 deployment, date-stamped 2026-10-06 against RunPod’s published rates (Community/Secure): H100 SXM $2.69/$3.49, H200 $3.59/$4.59, B200 $5.98/$6.79 per GPU-hour. Note these are multi-tenant “community” rates — dedicated-capacity contracts and sovereign-cloud premiums sit above them, and your negotiated number will differ.
One 2-node FP8 replica (16 H100, weights + KV headroom):
Community: $2.69 * 16 * 730h = $31,419/month per replica
Secure: $3.49 * 16 * 730h = $40,763/month per replica
One DeepSeek-shape disaggregated cell (4 prefill + 18 decode nodes):
Community: 22 nodes * 8 * $2.69 = $473/hr (~$346k/month)
Secure: 22 nodes * 8 * $3.49 = $614/hr (~$448k/month)
Hardware floor per 1M output tokens (18-node decode unit,
~14.8k out tok/s per node, DeepSeek's published node throughput):
RunPod H100: $0.40 / 1M output tokens
DeepSeek's own fleet at $2/hr H800 lease: $0.518 / 1M output tokens
The API comparison is the decision, not the architecture. Beam’s entire pitch is per-token efficiency — 23B active parameters means the marginal cost of a token on an efficiently packed fleet undercuts dense-500B models — but your fleet only realizes that efficiency at high utilization. At 30% utilization your effective cost triples. The honest planning number for a self-hosted Beam-class fleet is: build only if sustained utilization >50% or data-control requirements demand it; otherwise the API (Reflection’s own, or any Beam-hosting provider) is cheaper for the first year while the serving stacks mature.
When NOT to Build Any of This
- If your model is ≤70B-class dense or active. A single node holds it; PD-disaggregation and cells are over-engineering. vLLM’s own guidance: if the model fits on one GPU or one node, distributed inference is probably unnecessary.
- If your traffic is bursty and low-volume. Disaggregation amortizes over steady load; bursts wash out the prefill/decode split’s gains and you paid for a KV-transfer fabric you use twice a day.
- If you cannot budget the interconnect. Cross-node EP lives or dies by all-to-all latency. Without RDMA (or IB/NVLink across nodes), expert parallelism degrades to the “none” backend — all-reduce dispatch, which silently re-couples your throughput to the slowest link.
- If your team is <5 engineers. The DeepSeek pattern is 22 nodes per cell, three load balancers, an EPLB loop, and a KV-transfer fabric. That is a platform team’s worth of on-call surface for one model.
References & Further Reading
- Reflection: Introducing Beam (Oct 5, 2026) — the 501B/23B specs, RL-run scale, and efficiency claims.
- DeepSeek V3/R1 Inference System Overview — production parallelism shapes, load balancers, fleet stats, and the $87k/day cost line.
- DeepSeek-V3 Technical Report — the prefill/decode deployment sections and cross-node all-to-all design.
- SGLang: Expert Parallelism — backend selection, ep_size=tp_size constraint, DeepEP modes.
- SGLang: PD Disaggregation — the open-source replication path for the DeepSeek pattern.
- vLLM: Parallelism and Scaling — TP/PP placement guidance and the KV-cache size log lines to watch.
- vLLM: Data Parallel Deployment — DP-attention + EP for MoE models.
- NVIDIA Dynamo: Disaggregated Serving — the DGD spec and the when-not-to guidance.
- Slack Engineering: Cellular Architecture — the cell/drain pattern adapted for serving fleets.
- AWS Well-Architected: Cell-Based Architecture — the router/cell/control-plane decomposition.
- llm-d and AIBrix — Kubernetes-native serving stacks with prefix-aware routing.
- RunPod GPU pricing — rate card pulled live 2026-10-06.