Hub-and-Spoke vs. Mesh Cloud Networking: The Blast-Radius Argument Is Real, So Is the $1,460/Month Floor
Sources
- Hub-spoke network topology — Azure Cloud Adoption Framework
- Hub-spoke network architecture — Azure Architecture Center
- AWS Transit Gateway — Building a scalable and secure multi-VPC network infrastructure (AWS whitepaper)
- AWS Transit Gateway quotas — attachments, routes, bandwidth per AZ
- How transit gateways work — Availability Zone behavior (AWS docs)
- VPC peering basics — non-transitive routing (AWS docs)
- Amazon VPC quotas — routes per route table (500 default / 1,000 max)
- AWS Transit Gateway pricing (us-east-1, verified 2026-10-06)
- AWS Transit Gateway SLA — Multi-AZ 99.99% / Single-AZ 99.9%
- How Netflix is using IPv6 to enable hyperscale networking — AWS re:Invent 2021 NFX301
- Hyper-scale VPC Flow Logs enrichment — Netflix Tech Blog
- Scale generative AI use cases, Part 1: multi-tenant hub-and-spoke with AWS Transit Gateway — AWS ML Blog
- Migrate from hub-and-spoke to Azure Virtual WAN — Microsoft Learn
- Global transit network architecture and Virtual WAN — Microsoft Learn
- Azure Virtual WAN service guide — Azure Well-Architected Framework
- Azure Virtual Network pricing (verified 2026-10-06)
- VPC Network Peering — non-transitive routing (Google Cloud docs)
- Network Connectivity Center pricing — hub free, VPC spokes $0.10/hr, ADN $0.02/GiB
- Google Cloud VPC network pricing (inter-zone, inter-region)
Every cloud landing-zone engagement starts with a topology decision that is expensive to reverse: do you connect every network to every other network (mesh), or do you force traffic through a central hub? The vendor guidance is remarkably consistent — Azure's Cloud Adoption Framework, the AWS multi-VPC networking whitepaper, and Google's Network Connectivity Center all push hub-and-spoke — which is exactly why it deserves skepticism. When three competitors agree, the reason is usually that the managed hub is a revenue line for all of them.
So this guide takes the other side seriously. Mesh peering is genuinely cheaper for small N, genuinely lower-latency for directly-peered pairs, and genuinely less of a single point of failure. Hub-and-spoke wins on blast radius, route-table hygiene, and inspection — and Netflix's own re:Invent material documents what happens when you let a flat/mesh network grow past its limits. Both claims get receipts below: current quotas, current prices (verified against the live price lists on 2026-10-06), and the named deployments that prove each pattern survived production.
The two topologies, precisely
A mesh is N×(N−1)/2 direct connections with no transit hop. A hub-and-spoke is N attachments to one managed transit service. The important asymmetry: on every major cloud, direct peering is non-transitive. AWS states it plainly in the VPC peering docs: a VPC peering connection does not support transitive routing. Google's VPC Network Peering doc says the same: if net-a peers with net-b and net-a peers with net-c, net-b and net-c cannot communicate. Azure VNet peering has the same rule (peering is not transitive; you need user-defined routes or a hub NVA/firewall to chain spokes).
Non-transitivity is the load-bearing fact of the whole debate. It means "mesh" is not a scalable architecture — it is the absence of an architecture, and it degrades combinatorially as N grows:
Networks Mesh links Hub attachments
5 10 5
10 45 10
20 190 20
40 780 40
100 4,950 100
Mesh routes per spoke route table: N-1
Hub routes per spoke route table: 1 (default -> hub)At 40 networks, a full mesh needs 780 bidirectional peering links, and every spoke's route table carries 39 peer routes. AWS's route-table quota is 500 entries per route table by default, adjustable to a hard max of 1,000 with an explicit performance warning in the VPC quotas doc. Mesh route explosion and IP-space collisions (you must track every peer's CIDRs on both ends) are not edge cases at that scale — they are the operating model.
The named use case that killed mesh: Netflix at 100+ VPCs
The strongest public evidence for what breaks in flat/mesh networking is Netflix. Their re:Invent 2021 session NFX301, "How Netflix is using IPv6 to enable hyperscale networking", walks through the failure arc. Their requirements were "flat network, containers, continued growth, on-premises, 100+ VPCs, full IP reachability" — and the documented flaws of their existing IPv4 approach were:
- AWS routing limits ("VPC peering, etc.") as account and VPC count grew
- Private IPv4 address exhaustion, both in AWS and on-premises
- No security groups across the flat network boundary — the blast-radius problem stated bluntly, with an exclamation mark in the original deck
- Private IPv4 routing overhead and ENI density limits
The deck then shows the evolutionary ladder Netflix actually climbed for inter-VPC private reachability: customer gateway (VPN) → VPC peering → AWS Transit Gateway → internet gateway paths — before landing on IPv6 as the long-term flat-network substrate. That is the honest history: mesh peering was a rung, not a destination. Netflix's current production reality is a hybrid — their Cloud Network Insight post describes an ecosystem that mixes DirectConnect, VPC peering, Transit Gateways, and NAT gateways, with hundreds of thousands of VPC Flow Log files per hour. The lesson is not "hub always wins"; it is that pure mesh stops winning somewhere between tens and a hundred networks, and you should know which side of that line you are on before you pick.
A second named deployment worth knowing: AWS's own ML blog documents a multi-tenant hub-and-spoke pattern for generative AI — a centralized hub for shared AI service abstractions (model endpoints, vector stores) with tenant-specific spoke VPCs connected through Transit Gateway. That is the 2025–2026 shape of the pattern: not "hub for governance" but "hub for shared expensive things," which changes the cost math because AI east-west traffic between GPU fleets and data spokes is exactly the kind of traffic that pays transit data-processing fees twice if you design it wrong.
The architectures, provider by provider
AWS: Transit Gateway as the hub, peering as the mesh
AWS Transit Gateway is a regional, managed, multi-AZ transit service: one route domain that can attach VPCs, VPNs, Direct Connect gateways, and peered transit gateways. The verified capacity numbers from the quotas doc: up to 5,000 attachments per transit gateway (adjustable), 10,000 combined routes across its route tables, bandwidth up to 100 Gbps per VPC attachment per Availability Zone with up to 7.5 million packets per second per AZ, and an 8,500-byte MTU. The TGW SLA commits to 99.99% monthly uptime when deployed across two or more AZs (99.9% single-AZ).
Two operational facts that surprise teams migrating from peering meshes:
- Traffic enters the TGW only in the source AZ. Per how transit gateways work, if traffic is sourced from an AZ where the destination attachment is not present, the TGW internally routes it to a (random) AZ where the attachment exists — with no additional charge, but with a real, measurable cross-AZ hop. Latency-sensitive pairs (for example, a trading-ish API tier and its cache) should be peered directly or co-located in the same AZs, not forced through the hub.
- Route-table design is the actual blast-radius tool. One TGW supports up to 20 route tables (adjustable) with per-attachment association, so isolation domains are created by associating production/PCI/dev attachments to different route tables — segmentation without any mesh.
Azure: self-managed hub vs Virtual WAN
Azure gives you two hub implementations, and the choice is a control/scale trade. A self-managed hub-spoke is a VNet you build and operate: your own route tables, your own Azure Firewall or NVA, peering each spoke to the hub (intra-region peering, $0.01/GB each direction at both ends — verified from the VNet pricing page and the retail price API). Azure Virtual WAN is Microsoft-managed: the hub VNet, routing infrastructure, and gateways are operated by Microsoft, spokes attach via the same VNet peering, and any-to-any connectivity between spokes flows through the managed hub's routing fabric.
The named enterprise arc for Virtual WAN is the chemical distributor Azelis — “Azelis paints a fully connected global network with Azure Virtual WAN,” a Microsoft Azure case study with Wim Van de Water, Enterprise Cloud and Integration Architect, as the named source (the original customers.microsoft.com page has been retired; the study survives via CaseStudies.com's Microsoft Azure library). The pattern fits Azure's own Well-Architected service guide for Virtual WAN, which recommends the Standard SKU for production (zone redundancy, hub-to-hub and branch-to-branch connectivity) and sizes hub scale units against expected aggregate throughput. The global transit network architecture doc describes the model explicitly: any-to-any between globally distributed VNets, branches, and users — a hub-of-hubs (mesh of hubs) for planetary scale. The trade is real though: Virtual WAN enforces Microsoft's routing model, and the migration guide above exists precisely because moving an established self-managed hub to VWAN is not a config flip.
Google Cloud: peering-first, NCC when you need transit
GCP historically pushed VPC Network Peering (non-transitive, same rule as AWS) for the mesh case and let you build your own transit with Shared VPC or internal load balancers. The current hub product is Network Connectivity Center: a logical hub with VPC spokes, hybrid spokes (Interconnect/VPN), and router appliance spokes, providing transit between VPCs that could never reach each other through plain peering. Google's NCC pricing is structurally different from AWS/Azure: the hub itself is free, VPC spokes cost $0.10/hour each, and the Advanced Data Networking (ADN) data processing fee is $0.02/GiB for traffic that traverses a hub between VPC spokes — with the notable current waiver for hybrid spokes.
flowchart LR
subgraph MESH["Full mesh peering (N=6 shown)"]
A1[net a] --- A2[net b]
A1 --- A3[net c]
A1 --- A4[net d]
A1 --- A5[net e]
A1 --- A6[net f]
A2 --- A3
A2 --- A4
A2 --- A5
A2 --- A6
A3 --- A4
A3 --- A5
A3 --- A6
A4 --- A5
A4 --- A6
A5 --- A6
end
subgraph HUB["Hub-and-spoke"]
H[TGW / Virtual WAN hub / NCC hub]
S1[spoke 1] --- H
S2[spoke 2] --- H
S3[spoke 3] --- H
S4[spoke 4] --- H
S5[spoke 5] --- H
S6[spoke 6] --- H
endCost realism: the verified 2026 price sheet
All prices verified 2026-10-06 against the live AWS, Azure, and Google price lists and retail APIs — not against stale blog posts. AWS and Azure charge the hub both hourly (per attachment or per hub) and per-GB; Google charges spokes hourly and hub data processing per-GiB.
| Component | AWS (us-east-1) | Azure (eastus) | Google Cloud |
|---|---|---|---|
| Hub hourly floor | $0.05/hr per VPC attachment (TGW) | $0.25/hr Standard hub + $0.10/hr routing infra (VWAN) | Hub: $0. VPC spoke: $0.10/hr (NCC) |
| Hub data processing | $0.02/GB processed by TGW | $0.02/GB (VWAN hub data processed) | $0.02/GiB ADN through-hub |
| Mesh link cost | Peering: $0.01/GB in + $0.01/GB out (both ends) | VNet peering intra-region: $0.01/GB in + $0.01/GB out (both ends) | Peering: standard inter-zone rate $0.01/GiB (same-region); inter-region from $0.02/GiB |
| Central egress | NAT Gateway $0.045/hr + $0.045/GB | Azure Firewall Standard $1.25/hr + $0.016/GB | Cloud NAT (per-VM-hour, region-dependent) |
| Inter-region mesh | Inter-region VPC peering: $0.02/GB out, inbound free (US East to US West / EU Frankfurt verified); Asia-Pacific pairs higher ($0.09–$0.14/GB out) | Global VNet peering $0.035/GB in + $0.035/GB out (Zone 1; $0.09/$0.16 for Zone 2/3 pairs) | Inter-region VM-to-VM from $0.02/GiB (North America↔North America), $0.05–$0.14/GiB for other pairs |
Now the scenario that matters for sizing decisions: 40 spoke networks, 1 TB/month of east-west traffic between spokes.
AWS TGW hub-and-spoke:
40 attachments x $0.05/hr x 730 hr = $1,460.00
1 TiB data processing (1024 x $0.02) = $20.48
--------
= $1,480.48 / month
AWS full-mesh peering (same traffic):
780 links x $0 hourly = free
1 TiB x ($0.01 in + $0.01 out)/GB = $20.48
--------
= $20.48 / month
Hub premium at this size: ~ $1,460 / month
Azure VWAN (one hub):
hub $0.25 x 730 + routing $0.10 x 730 = $255.50
1 TiB hub data processing (1024x.02) = $20.48
(spoke peering to hub also applies) = $20.48
--------
= $296.46 / month
Google NCC:
hub: free
40 VPC spokes x $0.10 x 730 = $2,920.00
1 TiB ADN (1024 x $0.02) = $20.48
--------
= $2,940.48 / monthRead that carefully before you decree "hub everywhere" in your landing zone:
- The hub is a floor, not a metered convenience. AWS bills per attachment: 40 spokes cost $1,460/month even at zero traffic. GCP bills per spoke: $2,920/month at zero traffic for 40 VPC spokes (hub free, spokes billed). Azure's VWAN bills the hub (~$255/month for Standard + routing) plus peering — the cheapest floor of the three for pure spoke counts.
- Hub data processing is priced like a peering pair — until you inspect. AWS's pricing page is precise: data-processing charges apply “for each gigabyte sent from a VPC … to the AWS Transit Gateway” — once, on the sending side. Spoke-to-spoke through a TGW is $0.02/GB, the same $0.02/GB a peering pair costs ($0.01 out + $0.01 in). The fee only stacks when you insert a middle hop that sends again: route traffic through a firewall or appliance living behind its own TGW attachment and that hop's egress toward the TGW is billed as another “gigabyte sent from a VPC.” The inspection you get for it is the point; if you don't inspect, you're paying for a checkpoint with no guard.
- Hybrid is the honest default at scale. The cheap pattern is: hub for the long tail (onboarding a new spoke is one attachment, one route), direct peering for the 2–3 hot pairs that dominate billable GB (same-AZ path, no hub fee), and route-table isolation for the blast-radius domains. Netflix's production network is exactly this mix; so is most of the enterprise landing zone world.
Blast radius: the argument that actually decides it
Cost tables favor mesh at small N. Security favors the hub — and not for the reason the compliance slide says. The three real mechanisms:
- Choke points enable inspection. "All inter-VPC traffic passes through the hub" is only a useful property if something at the hub inspects: AWS Cloud WAN with Network Firewall, Azure Firewall in the VWAN secured hub (Standard: $1.25/hr + $0.016/GB, verified), or third-party NVAs in an NCC hub. Without the inspection appliance, hub-and-spoke is just a more expensive mesh with a single point of failure.
- Route tables are the isolation boundary. Non-transitive peering means a mesh has no central place to express "spoke 12 may not talk to spoke 38." Hub route tables (AWS: up to 20 per TGW; Azure VWAN: routing in the hub; GCP: spoke-attachment policies) make that a one-line route-table association instead of an audit of 780 peering links.
- Netflix's "no security groups" finding is the mesh endgame. Their deck lists the absence of a security-group enforcement boundary across the flat network as a top-level flaw of the IPv4 approach — blast radius in its purest form: any compromised workload had L3 reachability to the entire estate. A hub with per-attachment security-group/NACL enforcement and segmented route tables shrinks that to one routing domain.
When NOT to use hub-and-spoke
- N ≤ ~10 networks with stable pairs: 45 links max, route tables comfortably under quota, no compliance demand for centralized inspection — a hub is a $255–$1,460+/month floor solving a problem you don't have. Peer directly; revisit when N doubles.
- Latency-critical pairs: every hop through a regional hub adds AZ-hop latency (and TGW's documented random-AZ behavior means it is not always the nearest AZ). Colocate chatty tiers in the same VPC or AZ and peer directly — do not route a request path through a hub for governance aesthetics.
- Massive predictable east-west between two networks: if two spokes exchange tens of TB/month, the hub data-processing fee is pure tax with no inspection value on that path. Direct-peering that pair (a hybrid pattern, not a full mesh) is the standard fix.
- Single-region, single-team, one compliance boundary: the classic case for a landing zone with 3–5 VNets (prod/nonprod/shared). The hub's value — central onboarding, route segmentation, hybrid transit — materializes with team count and region count, not network count alone.
- Google Cloud at high spoke counts: NCC's per-spoke hourly (40 VPC spokes = $2,920/month floor) makes the full-hub pattern the most expensive of the three providers at this size. If your GCP estate is many projects with light traffic, Shared VPC (host project + service projects, no per-spoke hourly) is often the cheaper centralization — NCC earns its keep when you need hybrid transit or VPC-to-VPC transit that Shared VPC can't express.
And when NOT to use mesh
- N > ~20 and growing: you are managing N²/2 links and N×(N−1) routes against a 500-per-route-table default quota. Netflix's documented ceiling was reached in this regime.
- Any audit requirement for east-west inspection: mesh has no choke point; you cannot deploy one firewall for inter-VPC traffic. You can deploy 40 of them. That is a different conversation with your CFO.
- Overlapping CIDRs, ever: peering requires non-overlapping ranges on both ends; mesh multiplies the collision surface with every new network. A hub with centralized IPAM (or Cloud WAN/VWAN's address-space tooling) is the standard cure.
- Hybrid connectivity is in scope: on-prem must reach every VPC. In a mesh that is either N VPN tunnels or DirectConnect-to-every-VPC; in a hub it is one attachment. This alone decides most enterprise builds.
The decision tree
| Your situation | Honest answer |
|---|---|
| ≤10 networks, one team, no hybrid, no inspection mandate | Direct peering. Add the hub when N or auditors force it. |
| 10–50 networks, multiple teams, compliance asks "who can talk to whom" | Hub-and-spoke with segmented route tables. AWS TGW if you're AWS-heavy; Azure VWAN if Microsoft manages more of your estate than you want to admit; GCP: Shared VPC first, NCC when you need transit. |
| Global enterprise, many branches + many regions | Hub-of-hubs: VWAN global transit (any-to-any by design) or TGW inter-region peering. Mesh of hubs, never mesh of spokes. |
| 2–3 dominant heavy pairs inside a hub estate | Hybrid: hub for the tail, direct peering for the hot pairs. This is the Netflix production shape, not a compromise. |
| AI fleet with shared GPU/serving spoke + tenant data spokes | Hub with explicit capacity planning — the AWS multi-tenant hub-and-spoke AI pattern; watch the double-counted data-processing fee on model-traffic paths. |
References & Further Reading
- Hub-spoke network topology — Azure Cloud Adoption Framework
- Building a scalable and secure multi-VPC network infrastructure — AWS whitepaper
- How Netflix is using IPv6 to enable hyperscale networking — re:Invent 2021 NFX301 (PDF)
- Hyper-scale VPC Flow Logs enrichment — Netflix Tech Blog
- Multi-tenant hub-and-spoke for generative AI — AWS Machine Learning Blog
- Migrate to Azure Virtual WAN — Microsoft Learn
- Global transit network architecture and Virtual WAN — Microsoft Learn
- Azure Virtual WAN service guide — Azure Well-Architected Framework
- Network Connectivity Center overview — Google Cloud docs
- VPC Network Peering (non-transitive) — Google Cloud docs
- AWS Transit Gateway quotas · TGW pricing · TGW SLA
- Azelis: a fully connected global network with Azure Virtual WAN — Microsoft Azure case study (via CaseStudies.com)