Homa and the End of TCP for AI Clusters: Separating the Protocol From the Pitch

Sources

Every few years, someone announces the death of TCP in the data center. This time the announcement is aimed squarely at AI clusters, it comes from a person with the track record to be taken seriously, and — unlike most protocol proposals — it shipped as a working Linux kernel module you can install today without a reboot. John Ousterhout, Stanford professor emeritus and the person behind Raft, Tcl/Tk, and a shelf of distributed-systems papers, is now, in his own words, on a "life's mission" to replace TCP for data-center RPC with a protocol called Homa.

The why-now is real and checkable. His talk "Homa: The End of TCP for AI Clusters" landed at the AI Engineer World's Fair on September 17, 2026; The Register covered the campaign on October 1; the Hacker News thread ran October 4; and the Internet-Draft text in the homa-rfc repository was revised on October 3, 2026 to track the current implementation. The kernel module itself was pushed to the day we wrote this. This is not a dormant research artifact.

But "actively developed" and "you should run this in your GPU fleet" are different claims. This piece separates the two: what the protocol actually does, what the benchmark numbers do and do not establish, what is genuinely shipping versus what is still a README promise, what the GPU idle time behind TCP tail latency costs at 2026 prices, and — because no pattern is universally good — exactly when you should skip Homa entirely.

The Problem Is Real, Even If the Pitch Oversells It

The failure mode Ousterhout is attacking is not controversial. It is incast: several machines sending to one destination at once, overwhelming the last hop. If three incoming links each run at line rate toward a receiver whose single downlink drains at one line rate, packets queue up at the top-of-rack switch's egress queue. A tiny latency-sensitive message heading to that same receiver joins the back of that queue. Its own transmission time is microseconds; its wait time is dominated by the backlog in front of it.

  bulk sender A ───┐
                  │    ToR switch egress queue          receiver
  bulk sender B ───┼──► [ 40+ queued MTUs, low prio ] ──► downlink
                  │         ^                          (1x line rate)
  bulk sender C ───┘         │
                             │ trapped here behind bulk traffic
  short msg (KV probe) ──────┘   <- the message your GPU is waiting on

Two TCP properties make this worse. First, congestion control lives at the sender, far from the queue that is actually backing up. By the time an ECN mark or a packet loss travels receiver-side back to the senders and they adjust — a multi-round-trip convergence — the queue has already inflicted its latency, and the senders oscillate between overshooting and undershooting. Second, TCP is a byte stream: the transport cannot see message boundaries, so it cannot know that the 200-byte KV-cache probe behind a 900 MB checkpoint matters more, and cannot let it pass. That is head-of-line blocking in the transport, distinct from but compounding the switch-queue problem.

Why this bites AI fleets specifically: the traffic mix has changed. Training is still dominated by big, patient, throughput-bound transfers — and Ousterhout concedes outright that TCP and RDMA handle those fine. What is new is inference and agentic workloads: millisecond-scale compute cycles alternating with short coordination exchanges — distributed KV-cache lookups, worker barriers, metadata sync. When the compute phase itself is millisecond-scale, a millisecond-scale tail on the sync message is no longer noise. The barrier makes the whole group wait for the slowest exchange, and the GPUs sit idle for exactly that long.

The audience poll in the talk is worth mentioning precisely because of what it is not: an informal show of hands at a conference, with more hands than Ousterhout expected from people who suspect small-message latency is already limiting their applications. It is a diagnostic signal, not a measurement. Treat it that way.

How Homa Actually Works — and What Changed in September 2026

Homa's unit of communication is the RPC: an explicit-length request message and an explicit-length response, no connection setup. Three mechanisms replace TCP's sender-side guesswork:

Retransmission inverts TCP too: the receiver, which knows what it is missing, drives retransmits with RESEND packets rather than waiting for sender-side timeout archaeology. An anti-starvation FIFO escape hatch (the grant_fifo_fraction and pacer_fifo_fraction parameters) dedicates a slice of bandwidth to the oldest message so SRPT cannot starve a long transfer indefinitely. The protocol synopsis in the repository documents the full packet alphabet: DATA, GRANT, RESEND, UNKNOWN, BUSY, CUTOFFS, FREEZE, NEED_ACK, and ACK.

sequenceDiagram
    participant C as Client (GPU node)
    participant S as Server (KV cache / peer node)
    C->>S: START_MSG or unscheduled DATA (rtt_bytes)
    Note over S: message length now visible; SRPT ranks it against all inbound
    S->>C: GRANT (bytes + priority level)
    C->>S: scheduled DATA packets (paced)
    S->>C: response DATA (same grant mechanism in reverse)
    C->>S: ACK (server may now drop RPC state)

Note the diagram's first arrow: the protocol changed in September 2026, and most writing about Homa predates it. Per the module changelog: "Protocol change in granting mechanism. Messages that require grants are now completely scheduled (no initial unscheduled packets are sent). A new packet type, START_MSG, has been introduced for these messages to announce themselves to the recipient." That is a wire-format change — older clients and newer servers do not interoperate silently. Honest detail: the protocol.md synopsis in the same repository does not yet mention START_MSG, so the doc you will read first is one protocol revision behind the code you would run. Budget for that when evaluating.

Two more implementation realities from the same sources: rtt_bytes is a single global value — the synopsis notes that per-peer values for nonuniform-RTT fabrics are "not currently implemented," which matters if your fabric mixes near and far nodes; and the API is not sockets-compatible, so adopting Homa means application changes, not a sysctl. There is preliminary gRPC support and a Go client, which softens but does not eliminate that cost.

The Numbers — and What They Do Not Prove

The headline figure, from the benchmark in the talk and the USENIX ATC '21 kernel-implementation paper: on a 100 Gbps network at roughly 80% offered load, TCP's 99th-percentile round-trip for short messages exceeds 1.2 ms while Homa's sits at 92 microseconds — about 13x lower. Counterintuitively, the longest messages also improved, by nearly 2x, because run-to-completion scheduling beats TCP's fair sharing when the fabric is hot.

Now the caveats, which are load-bearing. The benchmark compares Homa against TCP — it does not quantify an advantage over RoCE/RDMA, which is what most serious AI fleets already run for their latency-sensitive paths. The workload is a synthetic request-response mix from the ATC paper (the repository's own cp_bench reproduces it), not a trace of a production inference fleet. And the talk's own curator states the application-level point plainly: the benchmark does not demonstrate a 13x end-to-end AI speedup. Throughput improves only to the extent that short-message waiting sits on the critical path — if sync is a rounding error next to a five-second compute phase, faster transport buys you nothing.

There is also a named, credible critic. Ivan Pepelnjak's January 2023 teardown of the position paper argues the ATC measurement setup biased the comparison: a small number of parallel TCP sessions carrying multiple independent messages, which manufactures transport-layer head-of-line blocking that production stacks — which use per-request connections or request multiplexing done properly — do not suffer the same way. His summary of the whole program: an attempt by a brilliant outsider to solve problems in someone else's field, where "marketing encounters reality." He is not reflexively anti-alternative — his own writing celebrates AWS's SRD — but his methodological objections have never been answered in print, and the 2026 talk did not address them. Neither do we dismiss them.

And one more straight from the horse's README: the SIGCOMM '18 paper's incast optimization — Section 3.6, the part of the academic design that specifically targets large incasts — has never been implemented in the module. The README asks anyone planning large-incast tests to contact the maintainer first. Hold that fact next to the pitch: the talk's central scenario is incast, and the module explicitly does not yet implement the paper's incast machinery. That does not make the shipping code useless — the grants-plus-priorities core is real and measured — but it is the single most important thing to know before a proof of concept.

What Is Actually Shipping

Verified against the repository and its changelog on October 5, 2026, here is the concrete state of the artifact:

Tooling exists and is honest: the repository ships cp_node and cp_bench for latency and throughput benchmarks, a /proc/net/homa_metrics interface exposing message-size distributions and counters, and sysctl knobs documented in the homa.7 man page. The maintainer explicitly invites pilot teams to contact him for setup help — worth taking him up on, because the alternative is discovering the unimplemented incast path in production.

The Incumbents Are Not Standing Still

The strongest argument against betting your fleet on Homa is that the hyperscalers already solved this problem their own way, in hardware, and you can rent their answer by the hour.

AWS built SRD (Scalable Reliable Datagram) into its fabric. The AWS Well-Architected Framework's performance pillar (PERF05-BP05) describes SRD as a transport that load-balances traffic across multiple paths and recovers quickly from packet drops — precisely the receiver-side, fabric-aware philosophy Homa pursues, implemented in the Elastic Fabric Adapter. Every EFA-backed EC2 instance — the Hopper-based P5 family and the Blackwell P6s — already runs a message-oriented, multipath, out-of-order-capable transport under your NCCL and MPI calls. If your fleet is on AWS, the honest question is not "Homa or TCP?" but "Homa or the transport I am already paying for?"

Google went hardware-first with Falcon and optical circuit switching. Google's Falcon transport — now a published Open Compute Project specification — is a NIC-resident, hardware-assisted transport built from production-proven components (Carousel, Snap, Swift, PLB, CSIG), carrying RDMA and NVMe as upper-layer protocols. And the TPU v4 paper describes the other half of the strategy: optically reconfigurable circuit switching that reshapes the physical topology to match each training job's communication pattern. When the network can physically re-route fibers for your all-reduce, a host-side transport protocol is attacking a smaller residual of the problem.

And the buy-side is rethinking the network too. Jane Street — an organization with zero tolerance for network hand-waving — put its networking leadership on tape discussing exactly this shift: how ML workloads are reshaping the networking layer, why the multicast that trading moved to decades ago (abandoning TCP for colocation-grade transport, as their exchange-architecture episode recounts) keeps failing to transfer to other domains, and what a new networking conference (Nines) says about the field's appetite for clean-slate designs. The Register separately reports Ousterhout is currently prototyping with one large financial-services company. Trading firms are the natural first adopters: they already run bespoke transport, own their switches, and price latency in dollars per microsecond.

TransportData modelCongestion control ownerWhere it runsAvailability to you
TCPByte streamSender (ECN/loss, multi-RTT convergence)EverywhereUniversal
RoCE v2 / RDMAMessage (RDMA ops)Sender + PFC/ECN fabric tuningNIC hardware + lossless fabricStandard on AI NICs (ConnectX, etc.)
AWS SRD (EFA)Message, multipath, out-of-order OKAWS fabric/NIC (receiver-side, per-path)AWS proprietary NICsRent it: every EFA instance (P4/P5/P6)
Google FalconConnection-oriented, request-responseNIC hardware (Swift-based, hardware-enforced shaping)Google fleet NICsOCP spec v1.1; not a product you install
HomaRPC messages, SRPT + grantsReceiver (grants) + switch prioritiesLinux kernel module, your switches as-isOpen source, out-of-tree, RHEL 8/9.5 backports

Read that last row against the two above it: Homa's genuine differentiator is that it is the only option in the table that runs on your existing commodity Ethernet and switches, with no NIC offload, no provider, and no lossless-fabric tuning. That is exactly the position SRD occupied inside AWS before AWS built hardware for it — and exactly why a self-hosted inference fleet on COTS hardware is Homa's real beachhead.

The Cost Math: What a Millisecond of GPU Idle Actually Is

The talk's core economic claim: even a millisecond of avoidable latency makes an expensive GPU sit idle, and GPUs are the most expensive idle silicon in the building. Let's price that honestly, with a date stamp — all prices verified October 5, 2026, and they will change.

A p5.48xlarge (8x NVIDIA H100, 3,200 Gbps) runs at $55.04/hour on-demand in us-east-1. That is $0.0153 per second — $0.0000153 per millisecond per node. Renting the same 8 accelerators via EC2 Capacity Blocks for ML lists at $5.970 per accelerator-hour as of the October 7, 2026 price update, i.e. $47.76/hour for the pod — capacity-block pricing has drifted below on-demand for P5, which tells you what the GPU market thinks of reserved-LLM-capacity risk right now.

Scenario (one p5.48xlarge pod, 8x H100)AssumptionAnnual cost of network-wait
Measured tail delta, barrier-heavy loop20 sync exchanges/sec, each paying the 1.2 ms TCP P99 vs 92 us Homa P99 from the benchmark~$10,700 per pod-year ($1.22/hr)
Same, 100-pod fleet (800 H100s)Same per-pod rate~$1.07M per fleet-year
Upper bound: 50% network-wait duty cyclems-scale cycles, half of each cycle spent waiting on sync (worst case in the talk's framing)~$241K per pod-year; ~$24.1M per 100-pod fleet-year

Two honesty notes on our own math. The 20-syncs-per-second figure is an illustrative model for an agentic loop with millisecond-scale cycles, not a measured production trace — your workload's barrier count and cycle time are the inputs that matter, and they swing the answer by orders of magnitude. And the 92 us vs 1.2 ms delta is Ousterhout's benchmark workload on a 100 Gbps fabric at 80% load; your fabric's load and topology will produce different tails. The point of the table is not the dollar figure — it is that the sign of the argument is right: at H100 pricing, tail latency on the coordination path converts directly into the most expensive idle time in your infrastructure, and it compounds across every pod and every barrier.

Against that, Homa's cost side is engineering time, not licenses: a kernel module to bake into your images (the repository's own utilities and our rollout estimate: roughly half a node-hour per node for image bake and canary rounds, so a 1,000-node fleet is on the order of 500 node-hours of toil before a single production flow moves), an application-side API migration for the RPC paths you move, and a wire-format change (September's START_MSG) that tells you protocol churn is still live. The qdisc work means your unmigrated TCP flows get faster too, which is the rare migration where the holdouts benefit.

When You Should Skip Homa

Skip it if your fleet is training-dominated. Ousterhout says it himself: for the big, patient, throughput-bound transfers that dominate LLM training, TCP and RDMA are fine. If your network's job is shuffling gradients and checkpoints and your barriers are rare and loose, there is no case — this is a protocol for mixed short-message traffic, not bulk transport.

Skip it if you are on a cloud provider's GPU fleet. You cannot install the module on p5 instances, and the provider's own transport — SRD over EFA on AWS, whatever the equivalent is inside Google's TPU fleet — already implements the same class of fix in hardware, with multipath and receiver-side awareness, maintained by people whose job it is. Your lever there is topology and placement, not host-side transport.

Skip it if you have not measured. Ousterhout's own closing recommendation is diagnostic before architectural: measure whether short-message P99 latency is actually limiting your application throughput. If sync wait is a small fraction of your cycle time, no transport protocol will buy you anything, and a kernel-module migration campaign is pure risk. The repository ships the measurement tooling (cp_bench, /proc/net/homa_metrics) — use it before committing to anything.

Proceed carefully if incast is your exact problem. The sharpest irony in the stack: the paper's Section 3.6 incast optimization is not implemented, so if your pain is specifically massive fan-in, you are pilot-testing the one scenario the README tells you to contact the maintainer about first. That is not disqualifying — the grants-plus-priorities core handles moderate incast, and the maintainer is responsive — but scope your PoC's claims to what the code does, not what the papers promise.

Consider it if you operate your own bare-metal inference or agentic fleet on commodity Ethernet, your cycles are millisecond-scale, your observability already shows sync paths at the top of your idle-time attribution, and you have switches with usable priority queues. That profile — self-hosted, latency-bound, COTS fabric — is where Homa has no incumbent competitor, and where a two-pod canary is cheap.

What We Would Do This Quarter

Not "adopt Homa." Do this instead, in order:

  1. Instrument the claim before the protocol. Add short-RPC P99 latency to your sync-path dashboards. If you cannot attribute GPU idle time to network waits today, no transport decision is grounded.
  2. Size the tail against your own duty cycle. Take your barrier rate and cycle time, and run the same per-millisecond math we did above against your actual per-node cost. The answer is either a rounding error or a line item — there is no middle.
  3. Exhaust the knobs you already own first. Per-message connections instead of multiplexed streams (Pepelnjak's critique cuts both ways: fix your usage pattern before blaming the transport), ECN thresholds on your ToR switches, and priority queues — the same primitive Homa leans on — configured by hand.
  4. If the numbers still hurt, run the canary. Two bare-metal nodes, the module's own cp_node benchmark against your real RPC mix, then one real KV-probe path. The maintainer answers his email; the repo tells you exactly which NICs are known-good. Budget a kernel-upgrade story for the out-of-tree module before it ever touches production.
  5. Watch three signals for the tipping point: the draft appearing in the IETF datatracker, mainline kernel acceptance (two years and counting), and any named production deployment with published numbers — the last one is the only proof that matters.

Our verdict: the problem is real and the design is principled, but the artifact is a research-grade kernel module with a live wire format, an unimplemented headline optimization, and a single maintainer's runway. That is exactly what a pilot-worthy protocol looks like two years before it is boring — and exactly what a production dependency does not look like today. Measure your tails, fix your TCP usage, and let the hyperscalers' hardware answers and Homa's kernel upstreaming race each other to your answer.

Related reading on the fleet side: our AI fleet architecture patterns, GPU capacity planning patterns, and Kubernetes GPU scheduling for LLM workloads cover the layers this protocol sits underneath, and LLM router gateways covers the request path above it.

References & Further Reading