Multi-Region Active-Active: The Patterns That Actually Ship (Netflix, Stripe, PlanetScale)
Sources
- Active-Active for Multi-Regional Resiliency — Netflix Tech Blog, Dec 2013
- Global Cloud — Active-Active and Beyond — Netflix Tech Blog, Apr 2017
- How Stripe's document databases supported 99.999% uptime with zero-downtime data migrations — Stripe, June 2024
- Stripe's DocDB: How Zero-Downtime Data Movement Powers Trillion-Dollar Payment Processing — InfoQ, Apr 2026
- REL13-BP02 Use defined recovery strategies — AWS Well-Architected Reliability Pillar
- DynamoDB global tables billing (rWRU mechanics) — DynamoDB Developer Guide
- DynamoDB global tables with multi-Region strong consistency now GA — AWS What's New, June 2025
- DynamoDB global tables — consistency modes, witness, FIS fault injection
- Amazon DynamoDB pricing (on-demand, us-east-1)
- PlanetScale read-only regions (pricing and replication semantics)
- AWS Fault Injection Service — DynamoDB global tables actions
- Zuul — Netflix edge gateway (GitHub)
- Denominator — Netflix multi-provider DNS control (GitHub)
Every architecture review with a compliance checkbox that says "multi-region" eventually produces the same slide: two regions, an arrow between them, and the word active-active in bold. The slide is fine. What the slide hides is that Netflix needed three years, a custom DNS control plane, and a purpose-built cache invalidation protocol to make active-active real — and that was for a workload (video streaming) whose data model is unusually forgiving. Stripe, whose data model is the least forgiving thing in the industry (money), still runs a single write region per shard set and spends its multi-region engineering on movement rather than multi-master writes.
This guide is the pattern catalog behind the slide: what active-active actually requires, the four patterns that have shipped at scale, what they cost at 2026 prices, and — most importantly — the decision tree for when active-passive with a fast data layer, or plain pilot light, is the honest answer.
First, the vocabulary: active-active is an RTO story, not a downtime story
The current AWS Well-Architected Reliability Pillar (REL13-BP02) orders multi-Region disaster recovery strategies by increasing cost and decreasing recovery objectives:
| Strategy | RPO | RTO | Recovery Region runs | The honest description |
|---|---|---|---|---|
| Backup and restore | Hours (PITR can reach ~5 min) | ≤ 24 hours | Nothing until disaster | You redeploy from IaC and restore data. Cheap, slow, rarely tested. |
| Pilot light | Minutes | Tens of minutes | Core data services only | Database and object storage are live and replicating; compute is a template. |
| Warm standby | Seconds | Minutes | A scaled-down full stack | Everything is running, small. Recovery is "scale up and shift traffic." |
| Multi-Region active-active | Near zero | Potentially zero | The full workload, both/all regions | No "recovery" event at all — but only if your data layer can accept writes in multiple regions without corrupting itself. |
Read the RPO column carefully. The pillar's own note is the most under-read sentence in the entire DR literature: "Data replication is useful for data synchronization and will protect you against some types of disaster, but it will not protect you against data corruption or destruction unless your solution also includes options for point-in-time recovery." An active-active database happily replicates your DELETE bug to every region in about half a second. That is not a footnote — it is the reason Stripe's write path is single-master and why we will spend most of this guide on the data tier.
The availability arithmetic that motivates the whole exercise, at 2026 prices:
99.9% -> 525.60 min/year (8.8 hours)
99.99% -> 52.56 min/year (under an hour)
99.999% -> 5.26 min/year (one bad deploy, twice a year
... maybe)
Nobody sustains five nines with a human in the loop. Five nines is not an on-call achievement; it is an architecture statement that no single region, and no single region's control plane, is allowed to be load-bearing. That is the actual product of active-active: not "zero downtime" marketing, but removing region and control-plane from your blast radius. Netflix's founding post says the same thing more bluntly: their driver was not AWS regions being flaky, but their own pace of change — "our pace of change sometimes breaks critical services in a region."
Pattern 1: Stateless active-active with async data replication (the Netflix pattern)
This is the canonical shape, shipped by Netflix in 2013 across US-East-1 and US-West-2 and extended globally in Global Cloud — Active-Active and Beyond (2017). The contract it imposes on the application is unforgiving and worth quoting as a checklist:
- Services must be stateless — all state lives in the data tier, which handles its own replication.
- All resources accessed on the user path must be local — if an S3 bucket is read on the user path, it must exist (and be populated) in every region; an app that publishes to S3 now publishes to N buckets.
- No cross-region calls on the user path — not "few". Zero. Any synchronous hop across regions puts the other region's fate back on your latency budget.
- Data replication must be asynchronous — synchronous cross-region writes couple the regions' availability; async keeps them independent, at the price of replication lag.
In steady state, users are geo-DNS routed to the closest region with a rough 50/50 split. On a regional failure, tools override geo-DNS and send everyone to the healthy region. Note what that implies for capacity: each region's fleet must be able to absorb ~100% of global traffic after the shift, or shed load deliberately. Netflix paired the DNS shift with Zuul edge-gateway traffic shaping to survive the thundering herd when a failed region's users reconnect all at once.
The data tier underneath was Apache Cassandra with multi-directional async replication, plus EVCache for caching and the custom CARP protocol to keep regional caches consistent with the source of truth. The number that made the whole design defensible: in Netflix's tests, a record written in one region was readable in the other within ~500ms. That is your actual consistency contract for this pattern — eventual, with sub-second typical convergence, and last-writer-wins when it isn't.
flowchart LR
U["Users (geo-resolved)"] --> GDNS["Geo-DNS / Route53\n~50/50 steady-state split"]
GDNS --> R1["Region A: full stack"]
GDNS --> R2["Region B: full stack"]
subgraph R1
Z1["Zuul edge gateway\n(traffic shaping, load shed)"] --> S1["Stateless services"]
S1 --> C1["Cassandra + EVCache\n(local reads/writes)"]
S1 --> B1["Regional S3 / SQS\n(fan-out writes x N)"]
end
subgraph R2
Z2["Zuul edge gateway"] --> S2["Stateless services"]
S2 --> C2["Cassandra + EVCache"]
S2 --> B2["Regional S3 / SQS"]
end
C1 <-. "async replication\n(typ. < 500ms convergence)" .-> C2
B1 <-. "async object replication" .-> B2
U -. "on region failure:\nDNS override to healthy region" .-> Z2
The DNS control plane deserves its own paragraph because it is where most homegrown attempts quietly die. Netflix deliberately rejected latency-based routing — "it could cause unpredictable traffic migration effects" — and used Denominator to drive a combination of UltraDNS (directional/geographic routing) and Route53 (fast, reliable config change), with the ability to push an emergency override. Latency-based routing during a partial failure is exactly when you do not want your DNS layer making autonomous traffic decisions; deterministic directional routing with a human (or runbook) override is far more predictable under stress.
What this pattern buys: no user-visible "recovery event" for the stateless tier, per-region isolation of every regional dependency, and a rehearsed traffic-shift runbook. What it costs: an application architecture review that outlaws cross-region synchronous anything, a cache-coherence protocol, and a data model that tolerates last-writer-wins on conflict.
Pattern 2: Single write region, replicated reads (the Stripe / PlanetScale pattern)
The second pattern drops the multi-master dream and keeps the latency win. One region owns writes for a given data set; other regions serve reads from asynchronously replicated copies. Two production reference points:
Stripe. The DocDB write-up (June 2024) and the follow-up InfoQ presentation (April 2026) describe a MongoDB-based, sharded DBaaS serving 5M+ queries/second across 2,000+ shards, with a fleet of Go proxies in front. Stripe's headline uptime target is 99.9995%. The architectural decision that matters: strong consistency for writes comes from a single point of authority per chunk, and multi-region resiliency comes from the ability to MOVE data and traffic, live — not from accepting writes in two places. Their Data Movement Platform migrates petabytes with client-transparent traffic switches: bulk import (sorted-order B-tree inserts, a 10x write-throughput win), bidirectional async replication from the oplog via Kafka + S3, a snapshot-diff correctness check, then a "versioned gating" traffic switch — bump a version token on the source shard so proxies refuse stale writes, let replication drain, flip the chunk route — completed in under two seconds. In 2023 they bin-packed 1.5 petabytes with this machinery and cut shard count by ~three quarters.
PlanetScale. The managed Vitess offering is explicit about the same bet: their read-only regions feature replicates your production keyspace asynchronously into remote regions for local reads, while writes stay in the primary. Replication lag is a first-class, queryable metric:
SELECT max_repl_lag();
-- instantaneous max seconds since the RO region
-- last stored a change made in the primary
This pattern is the honest default for most readers of this guide. It gets you local read latency everywhere, a much smaller blast radius than multi-master, and a rehearsed promotion path (promote the remote replica) — without ever needing conflict resolution. Its when-NOT-to-use is equally clear: if your writes are global and interactive (two users on two continents editing the same record in the same second), single-write-region adds a round-trip to every write for half your users, and no amount of read replication fixes that.
Pattern 3: Managed multi-master with strong consistency (the DynamoDB MRSC pattern)
The third pattern is new enough that most 2023-era architecture documents get it wrong. For a long time, "multi-master" in managed databases meant eventual consistency with last-writer-wins: DynamoDB global tables (MREC mode) replicate writes asynchronously, and concurrent writes to the same item in two regions resolve by timestamp — silent data loss by design. That was the trade: local write latency in every region, no cross-region read-your-writes.
Since general availability in June 2025, DynamoDB global tables can be configured for multi-Region strong consistency (MRSC): item changes are synchronously replicated so strongly consistent reads on ANY replica return the latest version, at the price of higher write latency than MREC. The design constraints are strict and worth knowing before you draw the happy path:
- Three replicas, or two replicas plus a witness. MRSC needs a quorum; the witness is a cost-optimized quorum member that stores replicated change data but serves no reads or writes — and per the billing documentation, replication to the witness incurs no rWRU, storage, or data-transfer charges.
- Empty-table conversion only. You cannot convert an existing single-region table with items to MRSC; the table must be empty during conversion.
- Same-account only, mode fixed at creation. MRSC global tables cannot mix consistency modes, and multi-account global tables remain MREC.
- Higher write latency than MREC. Synchronous cross-region replication puts the speed of light in your write path — measure before you commit.
# Step 1: create the base table with streams enabled
aws dynamodb create-table \
--table-name MusicTable \
--attribute-definitions \
AttributeName=Artist,AttributeType=S \
AttributeName=SongTitle,AttributeType=S \
--key-schema \
AttributeName=Artist,KeyType=HASH \
AttributeName=SongTitle,KeyType=RANGE \
--billing-mode PAY_PER_REQUEST \
--stream-specification StreamEnabled=true,StreamViewType=NEW_AND_OLD_IMAGES \
--region us-east-2
# Step 2: add a replica in a second region -> this table
# is now a global table (MultiRegionConsistency defaults
# to EVENTUAL; pass STRONG for MRSC on the conversion call)
aws dynamodb update-table --table-name MusicTable --cli-input-json \
'{
"ReplicaUpdates":
[
{
"Create": {
"RegionName": "us-east-1"
}
}
]
}' \
--region us-east-2
# Step 3: inspect replica state and consistency mode
aws dynamodb describe-table \
--table-name MusicTable \
--region us-east-2 \
--query 'Table.{TableName:TableName,MultiRegionConsistency:MultiRegionConsistency,Replicas:Replicas[*].{Region:RegionName,Status:ReplicaStatus}}'
The billing model is the part that surprises every finance review, because it is not "replication is free, you just pay more for the write." From the global tables billing page: a write to a replica is billed as one replicated write request unit (rWRU) in every region that has a replica table — and each GSIs you add doubles again (rWRU for the item plus a regular WRU for the GSI, in every region). The one genuinely generous clause: DynamoDB charges no cross-Region data transfer fees for global table replication — the bytes are free, the request units are not.
What it costs: order-of-magnitude math at 2026 prices
Every number in this section is scripted (Decimal, no mental arithmetic) against current DynamoDB on-demand pricing for us-east-1: $0.6250 per million write request units, $0.25/GB-month storage (Standard class), and PlanetScale's published read-only region rates. Prices change; re-run the math before you quote it.
writes/month: 259,200,000 (100/s x 30d)
write cost (us-east-1 rates):
1 region: $162.00/mo (259.2M WRU x $0.6250/M)
2 regions (global): $324.00/mo (2x rWRU billing)
3 regions (global): $486.00/mo (3x rWRU billing)
storage (1 TB table, Standard class):
1 region: $250.00/mo
2 regions: $500.00/mo (full copy in each region)
3 regions: $750.00/mo
PlanetScale read-only region (us-east-1):
PS-10RR replica: $16/mo
storage: $0.75/GB (100 GB -> $75/mo)
total for 100 GB: $91/mo per RO region
global-table write premium vs single region:
2-region: +$162.00/mo (+100%)
3-region: +$324.00/mo (+200%)
Read that premium line again, because it is the whole financial story of active-active in one row: going multi-region on DynamoDB does not add a "DR surcharge" to your writes — it multiplies your write bill by the number of regions. For write-heavy workloads the data tier cost scales linearly with region count, forever. The AWS pricing page's own worked example lands at $52.72 for 84.35M replicated write units in a 2-region table; our scripting reproduces the same product exactly.
Two more line items the slide never shows:
- Compute is the biggest bill, not the database. Active-active means the full fleet runs at full size in every region. A warm standby can run the recovery region at 30% and scale up in minutes; active-active cannot — that is the entire point. Double your compute, double your fleet-ops surface, and remember the Netflix capacity clause: each region must absorb ~100% of traffic after a DNS shift, so either overprovision steady state or engineer load shedding.
- Testing is a permanent line item. The current Reliability Pillar (REL13) is unusually blunt that untested DR is fiction, and DynamoDB now has first-class support: global tables integrate with AWS Fault Injection Service to pause replication between regions and rehearse region isolation in production-like conditions. If you are not running those experiments on a schedule, you do not have an active-active architecture — you have an assumption with a diagram.
The decision tree: which pattern do you actually need?
| Your situation | Honest pattern | Why |
|---|---|---|
| 99.9% target, single-region blast radius acceptable, tight budget | Backup + restore or pilot light | Hours of RTO is fine for most internal tools. Do not pay 2x compute for a compliance checkbox. |
| 99.99% target, regional read latency matters (global users) | Warm standby + read replicas / PlanetScale-style RO regions | Writes stay single-master (simple, safe), reads go local. The $91/mo PlanetScale example is the floor for a 100GB data set. |
| 99.99% target, data is the product, corruption is existential | Stripe pattern: single write authority + live data movement | Stripe targets 99.9995% uptime and still refuses multi-master writes. Money records do not tolerate last-writer-wins. |
| 99.999% target, writes are globally interactive, conflicts bounded by key | Managed multi-master with MRSC-style strong consistency | The June 2025 GA of MRSC makes this a real option where it previously meant accepting silent conflict loss. Three replicas or 2+witness, empty-table conversion, higher write latency. |
| 99.999% target, stateless services, data model tolerates eventual consistency | Netflix pattern: stateless active-active + async data replication | The only pattern with genuinely zero user-visible recovery events — if you can satisfy the four application contract rules (stateless, local resources, no cross-region calls, async replication). |
When NOT to do active-active
If your data model can't answer "who wins?" in one sentence, stop. Every multi-master system is a distributed-consistency problem wearing a marketing hat. MREC's answer is last-writer-wins (silent loss on conflict). MRSC's answer is a synchronous quorum (latency + complexity + empty-table conversion). If your team cannot state which of those answers applies to their tables, they are not ready to own either.
If you cannot afford to double the fleet, stop. The compute multiplier is non-negotiable in real active-active: every region runs everything. Warm standby exists precisely because most 99.99% workloads do not need the second full fleet until the day they need it.
If your writes are global and interactive, do not confuse read replication with a fix. Single-write-region means half your users pay an inter-region round trip on every write. If that latency is unacceptable AND strong consistency is required AND your conflict domain is per-key, MRSC is the first managed option that does not silently lose writes. If your conflict domain is "the row" in a relational sense, you are in Spanner/CockroachDB territory, with the operational maturity that implies — which most teams underestimate by roughly one team-year.
If nobody has rehearsed the failover, you have pilot light with extra steps. The REL13 pillar's anti-patterns list is explicit: leaving DR ad-hoc, depending on control-plane operations during recovery, and never testing the implementation are all "High" risk exposures. Netflix's geo-DNS override tooling and Stripe's two-second traffic switches were both built because a traffic shift you have not automated is a traffic shift you will improvise, badly, at 3 AM.
References & Further Reading
- Active-Active for Multi-Regional Resiliency — Netflix Tech Blog (the founding post: contract rules, Denominator, Zuul, Cassandra, EVCache)
- Global Cloud — Active-Active and Beyond — Netflix Tech Blog, 2017 (extending the pattern beyond the Americas)
- How Stripe's document databases supported 99.999% uptime with zero-downtime data migrations — Stripe Engineering (Data Movement Platform, versioned gating, <2s traffic switches)
- Stripe's DocDB: How Zero-Downtime Data Movement Powers Trillion-Dollar Payment Processing — InfoQ, April 2026
- REL13-BP02: Use defined recovery strategies — AWS Well-Architected Reliability Pillar (the current DR strategy ladder)
- Understanding Amazon DynamoDB billing for global tables — rWRU mechanics, no-transfer-fee clause, worked example
- DynamoDB global tables — MREC vs MRSC, witness, FIS fault injection, SLA tiers
- PlanetScale read-only regions — async RO replication, max_repl_lag(), pricing
- Zuul and Denominator — the Netflix edge and DNS control components, both open source
- AWS Fault Injection Service actions reference — region-isolation experiments for global tables