Model-Serving Topologies After Beam: When 501B Parameters Break Your Single-Cluster Instincts

Sources

Reflection shipped Beam on October 5, 2026: a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, trained on 23.8T tokens, with a reinforcement-learning run that burned 10,500 NVIDIA GB300 GPUs for four weeks to generate over 100 million rollouts. The weights are not public yet — Reflection promises the “weights, technical report, model card, and developer artifacts later this month.” That is the story. Not the benchmarks.

The benchmarks say Beam is competitive with GLM-5.2 while using 3–4× less inference compute, and it approaches Qwen 3.8-Max on coding and agentic tasks. Kimi K3 remains ahead on raw capability. Reflection’s own framing is honest here: the pitch is efficiency at inference time, not frontier supremacy. A 501B/23B sparse MoE at Western-open weights, arriving weeks after Kimi K3 and GLM-5.3, means the open-weight 500B class is now a routine purchase decision — and it breaks the single-cluster instincts most platform teams built around 70B-class models.

This guide is the serving-topology briefing for that world. The short version: a Beam-class model in FP8 “fits” on 8 H100s with about 11 GB left for KV cache. Fits is not serves. Everything below is about what actually serves.

The 501B/23B Trap: Fits-Is-Not-Serves

The arithmetic that catches teams is simple and merciless. Beam’s weights, by format:

BF16 (2 bytes/param):  501e9 * 2 = 1,002 GB across the fleet
FP8   (1 byte/param):  501e9 * 1 =   501 GB across the fleet

H100 80GB, BF16:  ceil(1002/80)     = 13 GPUs for weights alone
H100 80GB, FP8:   ceil(501/80)      =  7 GPUs for weights alone
H200 141GB, FP8:  ceil(501/141)     =  4 GPUs for weights alone

At 80% usable VRAM (weights + KV + activations + NCCL buffers):
H100 FP8:  ceil(501/(80*0.8)) = 8 GPUs, spare = 8*64-501 = 11 GB fleet-wide

Eleven gigabytes. A Beam-class GQA model with roughly 60 layers, 8 KV heads, 128-dim heads at FP8 KV eats ~120 KiB per token. Eleven GB buys 0.09M tokens of KV — at 32k context that is three concurrent requests across the entire replica, before activations. The “it fits” cluster is a machine that can hold the model and almost nothing else.

This is not a Beam defect. It is the defining property of the sparse-MoE class since DeepSeek-V3 (671B total / 37B active): weights memory scales with total parameters; compute per token scales with active parameters. The two axes are decoupled, and every topology decision below is a consequence of that decoupling. You buy memory for 501B and FLOPs for 23B. The 23B active side is genuinely cheap — Beam’s forward pass is ~46 GFLOPs/token, dense-23B arithmetic — which is why Reflection can claim GLM-5.2-class reasoning at a fraction of the compute. The 501B side is the bill.

The fix for the KV squeeze is never “more GPUs in the same replica.” Adding GPUs to a tensor-parallel replica shrinks per-GPU KV slightly (each rank holds a shard), but NCCL all-reduce communication grows and per-GPU KV still trails far behind what a serving workload needs. The fix is more replicas — which is where topology starts to matter.

Topology 1: The Aggregated Single-Cluster Default (And Its Ceiling)

The default topology every team starts with: one Kubernetes cluster, one vLLM or SGLang deployment, TP=8 per node, scale replicas horizontally under an HPA or llm-d-style gateway. It works because at 70B-class, a whole replica fits in one node’s NVLink domain and a single pod is the unit of both capacity and failure.

Beam-class models strain both properties at once:

The ceiling is not throughput. It is blast radius and utilization: a single-cluster fleet of expensive multi-node replicas is either sized for peak (idle money) or sized for average (queueing at peak). DeepSeek’s published fleet behavior shows the alternative.

Topology 2: Disaggregated Prefill/Decode — The DeepSeek Pattern

The most consequential production artifact for MoE serving is DeepSeek’s inference system overview (open-infra-index, Day 6). They run prefill and decode as separate deployments with different parallelism shapes, because the two phases have different bottlenecks:

Prefill unit:  4 nodes, 32 GPUs — routed-expert EP32, 9 routed experts
               + 1 shared expert per GPU, 32 redundant experts
Decode unit:   18 nodes, 144 GPUs — routed-expert EP144, 2 routed experts
               + 1 shared expert per GPU, 32 redundant experts

Why different: prefill is compute-bound (big batches, all-to-all hidden
               behind GEMMs); decode is memory-bound (small per-expert
               batches, all-to-all latency exposed on every token)

The decode unit is 4.5× the prefill unit in GPUs. That asymmetry is the whole point: prefill throughput scales with compute, decode throughput scales with KV and expert placement, and forcing them into one pool means one phase’s SLO is hostage to the other’s tail. Their numbers from a 24h production window (Feb 27–28, 2025): peak 278 nodes, average 226.75 nodes (8 H800/node), 608B input tokens (56.3% hitting the on-disk KV cache), 168B output tokens, ~73.7k input tok/s per node prefilling, ~14.8k output tok/s per node decoding, at an assumed $2/hr H800 lease = $87,072/day. Two load balancers — one for prefill (balances core-attention compute and dispatch send load), one for decode (balances KV usage and request counts) — plus an expert-parallel load balancer that detects high-load experts and re-replicates them every ~10 minutes. Nighttime, they shrink the inference fleet and hand nodes to research and training.

That daily flex is the pattern platform engineers should steal: inference capacity is a schedule, not a constant. Your 40% idle at 03:00 is someone else’s training capacity. DeepSeek’s own cost floor from those numbers is $0.518 per 1M output tokens at $2/hr — compare with your API bill before you build any of this.

The OSS replication path for this pattern: SGLang’s PD-disaggregation mode, or NVIDIA Dynamo’s DynamoGraphDeployment, whose disaggregated YAML is a three-line diff from aggregated:

services:
  Frontend:
    componentType: frontend
  prefill:
    componentType: worker
    subComponentType: prefill   # prompt processing only
  decode:
    componentType: worker
    subComponentType: decode    # token generation only

Dynamo’s docs are refreshingly blunt about when not to do this: “It is not automatically better. For small models, short prompts, low concurrency, or clusters without a fast KV-transfer fabric, an aggregated deployment is simpler and often faster.” The KV-transfer fabric is the hidden cost — prefill and decode pools only work if KV moves between them faster than it can be recomputed, which means RDMA or NVLink, which means network planning before the first pod lands.

Topology 3: Cell-Based Serving — Bounding the Blast Radius

The third topology is not about performance at all. It is about failure. A multi-node TP/EP replica is a shared-fate unit: lose one node and the replica’s KV cache is lost with it. Once your fleet is a handful of expensive replicas, the failure domain is your whole inference product.

The fix is the cell pattern, borrowed straight from general cloud architecture: Slack’s cellular migration made AZs drainable cells with an Envoy/xDS edge reweighting traffic; AWS’s Well-Architected cell guidance formalizes it: a thin router, N isolated cells, each cell handling a bounded slice of tenants. Applied to serving:

The cost is honesty about redundancy: N+1 cell redundancy (2 active + 1 warm standby) is a 50% hardware premium over N cells. For a 22-node cell fleet at RunPod Community rates, warm-standby redundancy adds $473/hr — about $346k/month. You buy that only when the product is enterprise-facing enough to need bounded blast radius, or when your cells are large enough that losing one is a page rather than an incident.

Picking a Topology: The Decision Table

TopologyReplica shapeBest fitCost driverWhen NOT to use
Aggregated single-clusterTP=8 per node, HPA replicas≤70B dense or small MoE; POCs; teams without RDMAGPU-hours × utilizationBeam-class MoE at FP8 floor (11GB KV headroom); multi-region SLOs
Disaggregated prefill/decode4-node prefill + 18-node decode units (DeepSeek shapes)Reasoning/agentic traffic (long prompts, long generations); ≥95% decode SLO needsKV-transfer fabric + node countShort prompts, low concurrency, no RDMA — aggregated is simpler and often faster (Dynamo docs)
Cell-based multi-cellN × (prefill+decode unit), thin routerEnterprise multi-tenant; bounded-blast-radius requirementsN+1 redundancy premium (~50%)<3 cells of demand — the router + control plane tax exceeds the resilience win

Read the table as cumulative, not exclusive: cells contain disaggregated units, which contain TP/EP replicas. The mistake is skipping levels — running cells without disaggregation wastes the cell boundary, running disaggregation without cells leaves the biggest failure domain untouched.

Cost Math That Survives a Finance Review

Order-of-magnitude economics for a Beam-class FP8 deployment, date-stamped 2026-10-06 against RunPod’s published rates (Community/Secure): H100 SXM $2.69/$3.49, H200 $3.59/$4.59, B200 $5.98/$6.79 per GPU-hour. Note these are multi-tenant “community” rates — dedicated-capacity contracts and sovereign-cloud premiums sit above them, and your negotiated number will differ.

One 2-node FP8 replica (16 H100, weights + KV headroom):
  Community: $2.69 * 16 * 730h  = $31,419/month per replica
  Secure:    $3.49 * 16 * 730h  = $40,763/month per replica

One DeepSeek-shape disaggregated cell (4 prefill + 18 decode nodes):
  Community: 22 nodes * 8 * $2.69 = $473/hr  (~$346k/month)
  Secure:    22 nodes * 8 * $3.49 = $614/hr  (~$448k/month)

Hardware floor per 1M output tokens (18-node decode unit,
  ~14.8k out tok/s per node, DeepSeek's published node throughput):
  RunPod H100:  $0.40 / 1M output tokens
  DeepSeek's own fleet at $2/hr H800 lease: $0.518 / 1M output tokens

The API comparison is the decision, not the architecture. Beam’s entire pitch is per-token efficiency — 23B active parameters means the marginal cost of a token on an efficiently packed fleet undercuts dense-500B models — but your fleet only realizes that efficiency at high utilization. At 30% utilization your effective cost triples. The honest planning number for a self-hosted Beam-class fleet is: build only if sustained utilization >50% or data-control requirements demand it; otherwise the API (Reflection’s own, or any Beam-hosting provider) is cheaper for the first year while the serving stacks mature.

When NOT to Build Any of This

References & Further Reading