Multi-Region Active-Active: The Patterns That Actually Ship (Netflix, Stripe, PlanetScale)

Sources

Every architecture review with a compliance checkbox that says "multi-region" eventually produces the same slide: two regions, an arrow between them, and the word active-active in bold. The slide is fine. What the slide hides is that Netflix needed three years, a custom DNS control plane, and a purpose-built cache invalidation protocol to make active-active real — and that was for a workload (video streaming) whose data model is unusually forgiving. Stripe, whose data model is the least forgiving thing in the industry (money), still runs a single write region per shard set and spends its multi-region engineering on movement rather than multi-master writes.

This guide is the pattern catalog behind the slide: what active-active actually requires, the four patterns that have shipped at scale, what they cost at 2026 prices, and — most importantly — the decision tree for when active-passive with a fast data layer, or plain pilot light, is the honest answer.

First, the vocabulary: active-active is an RTO story, not a downtime story

The current AWS Well-Architected Reliability Pillar (REL13-BP02) orders multi-Region disaster recovery strategies by increasing cost and decreasing recovery objectives:

StrategyRPORTORecovery Region runsThe honest description
Backup and restoreHours (PITR can reach ~5 min)≤ 24 hoursNothing until disasterYou redeploy from IaC and restore data. Cheap, slow, rarely tested.
Pilot lightMinutesTens of minutesCore data services onlyDatabase and object storage are live and replicating; compute is a template.
Warm standbySecondsMinutesA scaled-down full stackEverything is running, small. Recovery is "scale up and shift traffic."
Multi-Region active-activeNear zeroPotentially zeroThe full workload, both/all regionsNo "recovery" event at all — but only if your data layer can accept writes in multiple regions without corrupting itself.

Read the RPO column carefully. The pillar's own note is the most under-read sentence in the entire DR literature: "Data replication is useful for data synchronization and will protect you against some types of disaster, but it will not protect you against data corruption or destruction unless your solution also includes options for point-in-time recovery." An active-active database happily replicates your DELETE bug to every region in about half a second. That is not a footnote — it is the reason Stripe's write path is single-master and why we will spend most of this guide on the data tier.

The availability arithmetic that motivates the whole exercise, at 2026 prices:

99.9%   ->  525.60 min/year   (8.8 hours)
99.99%  ->   52.56 min/year   (under an hour)
99.999% ->    5.26 min/year   (one bad deploy, twice a year
                               ... maybe)

Nobody sustains five nines with a human in the loop. Five nines is not an on-call achievement; it is an architecture statement that no single region, and no single region's control plane, is allowed to be load-bearing. That is the actual product of active-active: not "zero downtime" marketing, but removing region and control-plane from your blast radius. Netflix's founding post says the same thing more bluntly: their driver was not AWS regions being flaky, but their own pace of change — "our pace of change sometimes breaks critical services in a region."

Pattern 1: Stateless active-active with async data replication (the Netflix pattern)

This is the canonical shape, shipped by Netflix in 2013 across US-East-1 and US-West-2 and extended globally in Global Cloud — Active-Active and Beyond (2017). The contract it imposes on the application is unforgiving and worth quoting as a checklist:

In steady state, users are geo-DNS routed to the closest region with a rough 50/50 split. On a regional failure, tools override geo-DNS and send everyone to the healthy region. Note what that implies for capacity: each region's fleet must be able to absorb ~100% of global traffic after the shift, or shed load deliberately. Netflix paired the DNS shift with Zuul edge-gateway traffic shaping to survive the thundering herd when a failed region's users reconnect all at once.

The data tier underneath was Apache Cassandra with multi-directional async replication, plus EVCache for caching and the custom CARP protocol to keep regional caches consistent with the source of truth. The number that made the whole design defensible: in Netflix's tests, a record written in one region was readable in the other within ~500ms. That is your actual consistency contract for this pattern — eventual, with sub-second typical convergence, and last-writer-wins when it isn't.

flowchart LR
    U["Users (geo-resolved)"] --> GDNS["Geo-DNS / Route53\n~50/50 steady-state split"]
    GDNS --> R1["Region A: full stack"]
    GDNS --> R2["Region B: full stack"]
    subgraph R1
        Z1["Zuul edge gateway\n(traffic shaping, load shed)"] --> S1["Stateless services"]
        S1 --> C1["Cassandra + EVCache\n(local reads/writes)"]
        S1 --> B1["Regional S3 / SQS\n(fan-out writes x N)"]
    end
    subgraph R2
        Z2["Zuul edge gateway"] --> S2["Stateless services"]
        S2 --> C2["Cassandra + EVCache"]
        S2 --> B2["Regional S3 / SQS"]
    end
    C1 <-. "async replication\n(typ. < 500ms convergence)" .-> C2
    B1 <-. "async object replication" .-> B2
    U -. "on region failure:\nDNS override to healthy region" .-> Z2

The DNS control plane deserves its own paragraph because it is where most homegrown attempts quietly die. Netflix deliberately rejected latency-based routing — "it could cause unpredictable traffic migration effects" — and used Denominator to drive a combination of UltraDNS (directional/geographic routing) and Route53 (fast, reliable config change), with the ability to push an emergency override. Latency-based routing during a partial failure is exactly when you do not want your DNS layer making autonomous traffic decisions; deterministic directional routing with a human (or runbook) override is far more predictable under stress.

What this pattern buys: no user-visible "recovery event" for the stateless tier, per-region isolation of every regional dependency, and a rehearsed traffic-shift runbook. What it costs: an application architecture review that outlaws cross-region synchronous anything, a cache-coherence protocol, and a data model that tolerates last-writer-wins on conflict.

Pattern 2: Single write region, replicated reads (the Stripe / PlanetScale pattern)

The second pattern drops the multi-master dream and keeps the latency win. One region owns writes for a given data set; other regions serve reads from asynchronously replicated copies. Two production reference points:

Stripe. The DocDB write-up (June 2024) and the follow-up InfoQ presentation (April 2026) describe a MongoDB-based, sharded DBaaS serving 5M+ queries/second across 2,000+ shards, with a fleet of Go proxies in front. Stripe's headline uptime target is 99.9995%. The architectural decision that matters: strong consistency for writes comes from a single point of authority per chunk, and multi-region resiliency comes from the ability to MOVE data and traffic, live — not from accepting writes in two places. Their Data Movement Platform migrates petabytes with client-transparent traffic switches: bulk import (sorted-order B-tree inserts, a 10x write-throughput win), bidirectional async replication from the oplog via Kafka + S3, a snapshot-diff correctness check, then a "versioned gating" traffic switch — bump a version token on the source shard so proxies refuse stale writes, let replication drain, flip the chunk route — completed in under two seconds. In 2023 they bin-packed 1.5 petabytes with this machinery and cut shard count by ~three quarters.

PlanetScale. The managed Vitess offering is explicit about the same bet: their read-only regions feature replicates your production keyspace asynchronously into remote regions for local reads, while writes stay in the primary. Replication lag is a first-class, queryable metric:

SELECT max_repl_lag();
-- instantaneous max seconds since the RO region
-- last stored a change made in the primary

This pattern is the honest default for most readers of this guide. It gets you local read latency everywhere, a much smaller blast radius than multi-master, and a rehearsed promotion path (promote the remote replica) — without ever needing conflict resolution. Its when-NOT-to-use is equally clear: if your writes are global and interactive (two users on two continents editing the same record in the same second), single-write-region adds a round-trip to every write for half your users, and no amount of read replication fixes that.

Pattern 3: Managed multi-master with strong consistency (the DynamoDB MRSC pattern)

The third pattern is new enough that most 2023-era architecture documents get it wrong. For a long time, "multi-master" in managed databases meant eventual consistency with last-writer-wins: DynamoDB global tables (MREC mode) replicate writes asynchronously, and concurrent writes to the same item in two regions resolve by timestamp — silent data loss by design. That was the trade: local write latency in every region, no cross-region read-your-writes.

Since general availability in June 2025, DynamoDB global tables can be configured for multi-Region strong consistency (MRSC): item changes are synchronously replicated so strongly consistent reads on ANY replica return the latest version, at the price of higher write latency than MREC. The design constraints are strict and worth knowing before you draw the happy path:

# Step 1: create the base table with streams enabled
aws dynamodb create-table \
    --table-name MusicTable \
    --attribute-definitions \
        AttributeName=Artist,AttributeType=S \
        AttributeName=SongTitle,AttributeType=S \
    --key-schema \
        AttributeName=Artist,KeyType=HASH \
        AttributeName=SongTitle,KeyType=RANGE \
    --billing-mode PAY_PER_REQUEST \
    --stream-specification StreamEnabled=true,StreamViewType=NEW_AND_OLD_IMAGES \
    --region us-east-2

# Step 2: add a replica in a second region -> this table
# is now a global table (MultiRegionConsistency defaults
# to EVENTUAL; pass STRONG for MRSC on the conversion call)
aws dynamodb update-table --table-name MusicTable --cli-input-json \
'{
  "ReplicaUpdates":
  [
    {
      "Create": {
        "RegionName": "us-east-1"
      }
    }
  ]
}' \
--region us-east-2

# Step 3: inspect replica state and consistency mode
aws dynamodb describe-table \
    --table-name MusicTable \
    --region us-east-2 \
    --query 'Table.{TableName:TableName,MultiRegionConsistency:MultiRegionConsistency,Replicas:Replicas[*].{Region:RegionName,Status:ReplicaStatus}}'

The billing model is the part that surprises every finance review, because it is not "replication is free, you just pay more for the write." From the global tables billing page: a write to a replica is billed as one replicated write request unit (rWRU) in every region that has a replica table — and each GSIs you add doubles again (rWRU for the item plus a regular WRU for the GSI, in every region). The one genuinely generous clause: DynamoDB charges no cross-Region data transfer fees for global table replication — the bytes are free, the request units are not.

What it costs: order-of-magnitude math at 2026 prices

Every number in this section is scripted (Decimal, no mental arithmetic) against current DynamoDB on-demand pricing for us-east-1: $0.6250 per million write request units, $0.25/GB-month storage (Standard class), and PlanetScale's published read-only region rates. Prices change; re-run the math before you quote it.

writes/month:            259,200,000  (100/s x 30d)

write cost (us-east-1 rates):
  1 region:              $162.00/mo   (259.2M WRU x $0.6250/M)
  2 regions (global):    $324.00/mo   (2x rWRU billing)
  3 regions (global):    $486.00/mo   (3x rWRU billing)

storage (1 TB table, Standard class):
  1 region:              $250.00/mo
  2 regions:              $500.00/mo   (full copy in each region)
  3 regions:              $750.00/mo

PlanetScale read-only region (us-east-1):
  PS-10RR replica:        $16/mo
  storage:                $0.75/GB     (100 GB -> $75/mo)
  total for 100 GB:       $91/mo per RO region

global-table write premium vs single region:
  2-region: +$162.00/mo  (+100%)
  3-region: +$324.00/mo  (+200%)

Read that premium line again, because it is the whole financial story of active-active in one row: going multi-region on DynamoDB does not add a "DR surcharge" to your writes — it multiplies your write bill by the number of regions. For write-heavy workloads the data tier cost scales linearly with region count, forever. The AWS pricing page's own worked example lands at $52.72 for 84.35M replicated write units in a 2-region table; our scripting reproduces the same product exactly.

Two more line items the slide never shows:

The decision tree: which pattern do you actually need?

Your situationHonest patternWhy
99.9% target, single-region blast radius acceptable, tight budgetBackup + restore or pilot lightHours of RTO is fine for most internal tools. Do not pay 2x compute for a compliance checkbox.
99.99% target, regional read latency matters (global users)Warm standby + read replicas / PlanetScale-style RO regionsWrites stay single-master (simple, safe), reads go local. The $91/mo PlanetScale example is the floor for a 100GB data set.
99.99% target, data is the product, corruption is existentialStripe pattern: single write authority + live data movementStripe targets 99.9995% uptime and still refuses multi-master writes. Money records do not tolerate last-writer-wins.
99.999% target, writes are globally interactive, conflicts bounded by keyManaged multi-master with MRSC-style strong consistencyThe June 2025 GA of MRSC makes this a real option where it previously meant accepting silent conflict loss. Three replicas or 2+witness, empty-table conversion, higher write latency.
99.999% target, stateless services, data model tolerates eventual consistencyNetflix pattern: stateless active-active + async data replicationThe only pattern with genuinely zero user-visible recovery events — if you can satisfy the four application contract rules (stateless, local resources, no cross-region calls, async replication).

When NOT to do active-active

If your data model can't answer "who wins?" in one sentence, stop. Every multi-master system is a distributed-consistency problem wearing a marketing hat. MREC's answer is last-writer-wins (silent loss on conflict). MRSC's answer is a synchronous quorum (latency + complexity + empty-table conversion). If your team cannot state which of those answers applies to their tables, they are not ready to own either.

If you cannot afford to double the fleet, stop. The compute multiplier is non-negotiable in real active-active: every region runs everything. Warm standby exists precisely because most 99.99% workloads do not need the second full fleet until the day they need it.

If your writes are global and interactive, do not confuse read replication with a fix. Single-write-region means half your users pay an inter-region round trip on every write. If that latency is unacceptable AND strong consistency is required AND your conflict domain is per-key, MRSC is the first managed option that does not silently lose writes. If your conflict domain is "the row" in a relational sense, you are in Spanner/CockroachDB territory, with the operational maturity that implies — which most teams underestimate by roughly one team-year.

If nobody has rehearsed the failover, you have pilot light with extra steps. The REL13 pillar's anti-patterns list is explicit: leaving DR ad-hoc, depending on control-plane operations during recovery, and never testing the implementation are all "High" risk exposures. Netflix's geo-DNS override tooling and Stripe's two-second traffic switches were both built because a traffic shift you have not automated is a traffic shift you will improvise, badly, at 3 AM.

References & Further Reading