GitOps Hub Scaling Limits: Where Argo CD and Sveltos Control Planes Actually Break
Hub-and-spoke GitOps won the architecture war, and now everybody is finding out what the hub costs. The question in every r/kubernetes thread about scaling a GitOps control plane is the same one: does the hub fall over at 5,000 applications, 20,000, or 50,000 — and is the fix more replicas, more shards, or a different controller? This guide answers it with published benchmark data, source-verified defaults, and runnable manifests for the two models worth comparing: Argo CD's replica-sharded controller and Sveltos's annotation-driven shard sets.
Two things make this worth writing down right now. First, Sveltos v1.16.0 shipped on October 5, 2026, overhauling pull-mode security with short-lived agent tokens and reduced RBAC — which changes the calculus for fleets where the management cluster cannot reach every spoke. Second, Argo CD v3.5.4 and v3.6.0-rc2 landed October 6, and the 3.6 security roll includes a repo-server memory-exhaustion advisory (GHSA-w996-f2wq-x9c6, oversized OCI manifests) plus a hard 4 MiB OCI-manifest rejection. When the security fixes of the week are about control-plane memory, scale is no longer a hypothetical.
The Structural Shift: The Hub Became the State Store
Both tools implement the same pattern: one management cluster holds a declarative picture of every spoke cluster, and controllers reconcile reality against it. The structural consequence is that the hub's working set grows linearly with fleet size, and everything it caches grows with it. A GitOps hub is not a stateless router you can put behind a load balancer. It is a cache of every managed resource in every cluster, plus a queue of work, plus (in Argo CD's case) a Redis instance that carries reconciliation chatter between components.
Where the two designs diverge — and where the scaling story actually lives — is what holds state and how you cut it apart:
- Argo CD keeps live state in three places: the Kubernetes API (Applications, ApplicationSets, AppProjects — all in etcd), Redis (manifest cache, revision metadata, and since v3.0 application health status), and in-memory cluster caches inside each application-controller replica. Scaling out means sharding the controller: you add replicas to the
argocd-application-controllerStatefulSet and must keep theARGOCD_CONTROLLER_REPLICASenvironment variable in sync, so each replica knows how many shards exist. Miss that sync and shards silently double-manage or orphan clusters. - Sveltos has no Redis and no separate API server. Every Sveltos resource — ClusterProfile, ClusterSummary, ClusterConfiguration — is a plain custom resource in the management cluster's etcd, and the management cluster's own Kubernetes API is the only entry point. Scaling out means cloning the whole controller set: an annotation on a managed cluster (
sharding.projectsveltos.io/key) makes a dedicated shard-controller deploy a parallel set of the five core controllers for that shard key. Remove the annotation and the shard set is torn down.
That difference sounds academic until you hit the wall. So let's put numbers on the wall first.
Where the Walls Actually Are: The Benchmark Evidence
The most complete public scaling data for Argo CD is the CNOE benchmarking study (Andrew Lee and Gaurav Dhamija of AWS, with Michael Crenshaw of Intuit) — 500 workload clusters and 50,000 applications, the largest published Argo CD scale test. Be honest about its vintage when you use it: it was run against Argo CD 2.8 in late 2023, and every workload was a 2 KB ConfigMap. Your Helm-rendered Deployments are fatter, and v3.x changed health-status storage into Redis. Use the shapes, not the absolute minutes.
Four results from that study decide most hub-sizing arguments:
| Experiment | Setup | Result |
|---|---|---|
| Client QPS / burst | 15,000 apps, default QPS 50/burst 100 vs 100/200 | Sync time fell from 61.5 min to 29.5 min — a 52% cut from one config change. Gains flatten past QPS 150/burst 300 (~19 min) and vanish after ~250/500. |
| Cluster scaling | 10 shards, one m5.2xlarge (8 vCPU / 30 GB), 10k apps | 100 clusters → 9 min sync; 250 → 9 min; 500 → 11 min. The wall was the node, not the controller. |
| Application scaling | 10 shards, 500 clusters, adding m5.2xlarge nodes | 15k apps saturated one node at 100% CPU/memory. 20k and 25k synced in ~11–12 min on two nodes; 30k hit the two-node ceiling (19 min, metric blips); 50k needed a third node and took 22 min. |
| Shard count | 500 clusters, 50k apps, 3 nodes | 3 shards → 75 min; 6 → 37 min; 9 → 21 min. Past 9 shards, zero improvement — the nodes were CPU-saturated. Shards only help if there is idle CPU to consume. |
The two lessons that survive every version bump: the first scaling lever is configuration, not hardware (client QPS alone halved sync time), and sharding buys throughput only while the underlying nodes have spare CPU. A shard on a full node is a process, not a win.
The rebalancing tax nobody budgets for
The same study's sharding deep dive (100 clusters, 10k apps, 10 shards) measured what happens to cluster-to-shard assignments when you shrink 10 shards to 9 — routine during node maintenance or autoscaler activity:
| Sharding algorithm | CPU variability across shards | Memory variability | Assignment changes at 10→9 shards |
|---|---|---|---|
| legacy (uid-hash) | 0.32 (0.23–0.55 core spread) | 227 MiB | n/a |
| round-robin (uniform clusters) | 0.02 | 110 MiB | n/a |
| round-robin (random apps per cluster) | 0.27 | 136 MiB | n/a |
| greedy minimum (by app count) | 0.06 | 109 MiB | 75 of 100 clusters moved |
| weighted ring hash | 0.12 | 163 MiB | 48 moved |
| consistent hashing with bounded loads | 0.17 | 131 MiB | 15 moved |
Read that last column carefully. Greedy minimum balances load beautifully, but a one-shard change reassigned three quarters of the fleet — and every reassignment means a new controller replica cold-building its cluster cache: watch bursts, API-server load, and a temporary memory spike on every affected shard. Consistent hashing with bounded loads moved only 15 clusters at the cost of slightly worse balance. That is exactly the trade the Argo CD team made when they shipped consistent-hashing as the third --sharding-method option — still labeled alpha in the high-availability docs, with round-robin known to reshuffle all clusters when the rank-0 shard is removed. Treat both algorithms as production-tested-by-others, not production-proven.
Sveltos sidesteps the reshuffle tax differently: shard membership is not computed, it is declared in the cluster's annotation, so scaling a shard out never redistributes another shard's clusters. The cost is that you own the sharding policy — a badly skewed annotation plan gives you a hot shard and no algorithm to fix it.
Architectural Blueprint: Two Ways to Cut a Hub
The same fleet, drawn both ways. Note where state lives in each — that is where the memory goes and where the blast radius starts.
ARGO CD: single controller, replica-sharded SVELTOS: controller SETS, annotation-sharded
mgmt cluster mgmt cluster
┌────────────────────────────────┐ ┌─────────────────────────────────────┐
│ argocd-server │ │ K8s API (the ONLY entry point) │
│ repo-server ──fork/exec──► │ │ ┌──────────────┐ ┌──────────────┐ │
│ Helm/Kustomize renders │ │ │ shard-set A │ │ shard-set B │ │
│ Redis ◄── manifest cache, │ │ │ addon-ctrl │ │ addon-ctrl │ │
│ revision info, │ │ │ classifier │ │ classifier │ │
│ health status (v3.0+) │ │ │ event-mgr │ │ event-mgr │ │
│ app-controller StatefulSet: │ │ │ sc-mgr │ │ sc-mgr │ │
│ replica0 shard0 ──┐ │ │ │ hc-mgr │ │ hc-mgr │ │
│ replica1 shard1 │ in-mem │ │ └──────┬───────┘ └──────┬───────┘ │
│ replica2 shard2 ──┘ cluster │ │ │ watched by │ │
│ caches │ │ shard-controller: on new annotation
│ state: etcd (Apps) + Redis │ │ sharding.projectsveltos.io/key on a
│ + per-replica RAM │ │ cluster → deploy a new controller SET
└──────────┬───────────────────┘ │ state: etcd only (CRs). NO Redis.
│ watch/poll every spoke └──────┬──────────┬──────────────┘
spoke1 ───┴─── spoke2 ──── spoke3 spoke1 spoke2
(no agent; hub needs outbound (sveltos-agent runs IN each spoke in
access to spokes) agent mode; pull mode + SHORT-LIVED
tokens since v1.16.0)Three consequences fall straight out of this diagram:
- Redis is Argo CD's single fastest-moving failure domain. Lose it and you lose the manifest cache, revision metadata, and (since v3.0) health status — the controllers keep running but every app reconciles as a cache miss, which turns a Redis blip into a fleet-wide repo-server CPU storm. Sveltos has no equivalent component to lose; its "cache" is etcd, which you are running anyway.
- Argo CD needs outbound reach from hub to spoke; Sveltos offers pull mode. If your spokes sit behind NAT or in air-gapped networks, Sveltos v1.16.0's short-lived-token pull mode is an architectural fit, not a workaround. With Argo CD you'll be engineering connectivity (VPNs, tunnels) for the hub.
- The sharding unit is different, so the ops model is different. Adding Argo CD capacity = edit StatefulSet replicas + keep one env var in sync (or adopt the alpha dynamic distribution). Adding Sveltos capacity = annotate clusters, let the shard-controller deploy a whole parallel stack, and tune each set through the shard-components ConfigMap.
Production Manifests: Scaling a Hub Without Lying to Yourself
Everything below uses verified defaults: Argo CD values were read from the v3.5.4 source tree (install manifests and controller/repo-server main files), Sveltos values from its current docs. Nothing here is guessed from a blog post.
1. Argo CD: shard the controller without the env-var foot-gun
The classic failure is setting StatefulSet replicas: 3 and forgetting ARGOCD_CONTROLLER_REPLICAS — shard math silently miscounts. This Kustomize patch does both, pins the alpha-but-benchmarked sharding algorithm, and raises client QPS (the single highest-ROI change in the CNOE data):
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: argocd-application-controller
spec:
replicas: 3 # shard count = replica count. Keep both numbers equal.
template:
spec:
containers:
- name: argocd-application-controller
env:
- name: ARGOCD_CONTROLLER_REPLICAS
value: "3" # MUST equal spec.replicas — sharding miscounts otherwise
- name: GOMEMLIMIT
value: "1800MiB" # ~90% of a 2Gi limit: GC fires before the OOM killer does
# Client QPS against spoke API servers. Defaults in v3.5.4 source: QPS 50,
# burst = 2x QPS (100). CNOE: raising these halved sync time in one change.
- name: ARGOCD_K8S_CLIENT_QPS
value: "100"
- name: ARGOCD_K8S_CLIENT_BURST
value: "200"
resources:
requests:
cpu: "1"
memory: "1Gi"
limits:
memory: "2Gi" # the in-mem cluster cache is the OOM victim; size for fleet
---
apiVersion: v1
kind: ConfigMap
metadata:
name: argocd-cmd-params-cm
data:
# Sharding algorithm: legacy (uneven), round-robin (even, reshuffles on rank-0 loss),
# consistent-hashing (bounded loads; 5x fewer reassignments in CNOE tests). Alpha status.
controller.sharding.algorithm: "consistent-hashing"Two warnings with this patch. First, the QPS variables above throttle the controller's watch/refresh traffic against spokes; check your hub's own API-server Priority and Fairness queues before you go further — starving system-level flows on a shared hub cluster is a self-inflicted outage. Second, if you change shard counts often, look at the dynamic cluster distribution feature (alpha since v2.9): it moves shard count from the env var to the Deployment's replica field, so scaling no longer restarts every controller pod, and a heartbeat ConfigMap tracks which pod owns which shard.
2. Argo CD: the repo-server and reconciliation knobs people forget
The repo-server renders every manifest with fork/exec Helm and Kustomize processes. Its --parallelismlimit defaults to 0 — no limit at all in v3.5.4, which means one bursty monorepo webhook can fork enough concurrent renders to OOM the pod (the docs recommend capping it exactly for this reason). Pair the cap with jitter so thousands of apps don't refresh in lockstep:
apiVersion: apps/v1
kind: Deployment
metadata:
name: argocd-repo-server
spec:
template:
spec:
containers:
- name: argocd-repo-server
args:
- --parallelismlimit
- "8" # default is 0 = UNLIMITED fork/exec renders. Cap it.
resources:
requests:
memory: "512Mi"
limits:
memory: "1Gi"
env:
- name: GOMEMLIMIT
value: "900MiB" # same OOM-mitigation trick as the controller
---
apiVersion: v1
kind: ConfigMap
metadata:
name: argocd-cm
data:
# Git poll default is 3m. Jitter (default 60s) stops a mass-refresh thundering herd.
timeout.reconciliation: "180s"
timeout.reconciliation.jitter: "60s"
# Webhook stampede control: when one monorepo push affects >10 apps,
# spread their refreshes over 60s instead of one queue spike.
webhook.refresh.jitter: "60s"
webhook.refresh.jitter.threshold: "10"
# Plain-YAML monorepos: disable webhook manifest-cache warming entirely
# (recommended in the HA guide for exactly this shape of repo).If your repos are plain YAML (no Helm/Kustomize), also set per-repo webhookManifestCacheWarmDisabled via the CLI (argocd repo edit) — the HA guide explicitly recommends it for large monorepos, because warming caches for thousands of unaffected apps on every push buys nothing and slows the webhook handler's Redis round-trips.
3. Sveltos: scale out with an annotation, not a replica dance
Sveltos sharding is opt-in per cluster group. Annotate a set of managed clusters with the same shard key and the shard-controller deploys a dedicated five-controller stack for them; tune that stack through one ConfigMap that supports Strategic Merge and JSON (RFC 6902) patches, targeted by component-name regex:
# 1. Tag a cluster into shard "eu-west". The shard-controller does the rest.
apiVersion: lib.projectsveltos.io/v1beta1
kind: SveltosCluster
metadata:
name: spoke-eu-1
annotations:
sharding.projectsveltos.io/key: "eu-west" # same key = same controller set
---
# 2. Per-shard-set tuning, delivered via --shard-components-config on the shard-controller.
# Each data key is one independent patch. Target by group/kind/name-regex.
apiVersion: v1
kind: ConfigMap
metadata:
name: shard-components-config
namespace: projectsveltos
data:
node-scheduling: |-
patch: |-
apiVersion: apps/v1
kind: Deployment
metadata:
name: any-name # required by SMP; the target field does the matching
spec:
template:
spec:
tolerations:
- key: dedicated
operator: Equal
value: sveltos-shards
effect: NoSchedule
nodeSelector:
node-role: sveltos-shards
target:
group: apps
kind: Deployment # applies to ALL five components in the shard set
healthcheck-memory: |-
# hc-mgr polls every managed cluster and runs hotter than the rest:
# target ONLY it. Name is an anchored regex; hc-manager.* matches
# hc-manager-eu-west (shard key is part of the deployment name).
patch:
- op: add
path: /spec/template/spec/containers/0/resources
value:
requests:
memory: "512Mi"
target:
group: apps
kind: Deployment
name: hc-manager.*Before you shard at all, tune the single set. The two knobs live in the addon-controller source as --concurrent-reconciles (default 10) and --worker-number (default 20), and the docs are explicit about the threshold: raise worker-number when managed clusters exceed ~100. In the official chart, pass them through the addon-controller's extraArgs map (verified against the chart's values.yaml at v1.16.0):
# projectsveltos chart (helm upgrade --install projectsveltos projectsveltos/projectsveltos)
# Source of flag names + defaults: addon-controller pkg/app/app.go at v1.16.0
# --concurrent-reconciles default 10 (parallelism for ClusterProfile/ClusterSummary)
# --worker-number default 20 (workers for long tasks: Helm installs, kustomize builds)
addonController:
controller:
extraArgs:
concurrent-reconciles: 20
worker-number: 40And since v1.16.0, fleets whose spokes cannot be reached from the hub should register in pull mode: the spoke-side agent initiates contact and now authenticates with short-lived tokens instead of a long-lived credential stored on the spoke — a materially smaller blast radius if a spoke is compromised, and the RBAC the agent gets is limited to what it actually needs.
Day-2: What to Watch, and What Breaks First
Metrics that earn their dashboard panels
Argo CD exposes per-component Prometheus metrics; these are the ones that move before an incident (names verified against the v3.5.4 metrics documentation):
# 1. Reconciliation latency is creeping up (the earliest wall signal)
histogram_quantile(0.95,
sum by (le) (rate(argocd_app_reconcile[15m])))
# 2. Redis trouble: failed commands during reconciliation (cache misses are
# expected on cold start; sustained non-zero after warm-up is a Redis problem)
sum by (command) (rate(argocd_redis_request_total{failed="true"}[10m]))
# 3. Cluster cache freshness — a stalled cache means stale drift detection
max(argocd_cluster_cache_age_seconds)
# 4. API pressure against spokes / retries climbing = your QPS is wrong
sum(rate(argocd_app_k8s_request_total[5m]))
sum(rate(argocd_kubectl_request_retries_total[5m]))Sveltos, having no Redis, surfaces health through its own dashboard and Grafana stack plus plain controller workqueue metrics — and v1.16.0 added agent-side error reporting, so classification and health-check failures on a spoke now show up as status instead of silent log-spelunking.
Runbook notes from the field-shaped failure modes
- The OOMKilled cold-start loop. A controller pod that dies and re-syncs its cluster cache will often OOM again, because cache rebuild is the memory peak. The documented mitigation is GOMEMLIMIT at 80–90% of the container limit — set too close to the working set it causes GC thrashing instead, so if you see high CPU plus frequent GC while memory sits pinned, raise the limit rather than tune the percentage.
- "continue parameter is too old." On clusters with very large resource counts, controller cache syncs can outrun the etcd compaction window and fail with this error. The documented fix is
ARGOCD_CLUSTER_CACHE_LIST_PAGE_BUFFER_SIZE(withARGOCD_CLUSTER_CACHE_LIST_PAGE_SIZE), which pre-fetches list pages into memory — note the trade: memory for reliability, so re-check your limits after raising it. - The reshuffle storm. On round-robin sharding, removing the rank-0 cluster reshuffles every cluster across shards at once. If you must remove a shard, do it during low change traffic and expect a cache-rebuild spike; consistent-hashing exists precisely to shrink this from "all clusters" to "a handful."
- Rate-limiter defaults are asleep. The controller's workqueue global limiter is disabled by default (bucket 500, QPS unlimited) and per-item exponential backoff is off (
WORKQUEUE_FAILURE_COOLDOWN_NS=0). A misbehaving app that requeues constantly can peg the controller until you setWORKQUEUE_BUCKET_QPS. Know these exist before you need them. - Spoke churn is the hidden killer for AI/ML fleets. The CNOE study called this out explicitly: ephemeral clusters spinning up and down trigger constant shard reassignment under app-count-based algorithms. If your fleet is ephemeral (training clusters, CI runners), stable-shard schemes beat optimally-balanced ones.
Blast radius, honestly stated
| Failure | Argo CD blast radius | Sveltos blast radius |
|---|---|---|
| One controller shard/replica dies | Its clusters go unmanaged until reschedule + cache rebuild (minutes). Other shards unaffected. | Same — one shard set's clusters unmanaged. Others fine. |
| Redis dies (Argo CD only) | Whole hub degrades: every reconcile becomes a cache miss → repo-server render storm, UI slows, health status (stored in Redis since v3.0) goes stale. | Component does not exist. |
| etcd / management API down | Total hub outage; spokes keep running last-applied state. | Total outage; same running-state resilience on spokes. |
| Hub→spoke network partition | Spokes drift silently; hub queues work and hammers retries. | Same — unless spokes are registered in pull mode, in which case agents keep their last instructions and re-establish with short-lived tokens. |
| Shard scale-down mistake | Rank-0 removal reshuffles all clusters (round-robin); consistent-hashing moves few. | Annotation removal only affects that shard's clusters — but a wrong annotation plan strands a cluster with no controller. |
That middle row is the honest one-line summary of the whole debate: the spoke clusters do not stop running when the hub breaks — they stop being correct. Size the hub for your tolerance of undetected drift, not for your app count.
The Decision Matrix
| Your fleet | What breaks first | Do this |
|---|---|---|
| < ~500 apps, < 20 clusters | Nothing. Genuinely. | Defaults are fine. Don't shard — you'd be adopting an ops burden (and alpha algorithms) for zero throughput. |
| 500–5,000 apps | Sync latency during bursts; webhook stampedes. | QPS/burst tuning (biggest lever in the CNOE data), reconciliation + webhook jitter, repo-server parallelism cap. Still no shards. |
| 5,000–15,000 apps | Repo-server CPU/memory on render bursts; controller RAM growth. | Set resource limits + GOMEMLIMIT, split monorepos or disable cache warming, watch argocd_app_reconcile p95. First shard conversation. |
| 15,000+ apps or 100+ clusters (Argo CD) | The node: one m5.2xlarge saturated at 15k apps in the benchmark. | Shard properly (replicas + env var + algorithm choice), add nodes with shards, accept the reshuffle tax, or evaluate whether the app-level granularity is the problem. |
| Spokes unreachable from hub / air-gapped / ephemeral | Connectivity engineering, not controller CPU. | This is Sveltos's home turf: pull mode with v1.16.0 short-lived tokens, agentless or agent modes, annotation sharding for fleet subsets. |
| Add-on fleet management (CNI, cert-manager, observability agents) rather than app delivery | Misusing an app-delivery engine for fleet day-0/1 concerns. | Sveltos ClusterProfiles target clusters by label selector and deploy Helm/YAML/Kustomize per cluster — purpose-built for the "same add-on, 200 clusters" job Argo CD does awkwardly. |
One more honest note: if operating any of this is more than your team wants to own, the commercial path exists — Akuity (founded by Argo CD's creators) sells a managed enterprise delivery platform and advertises control over a million-plus applications; Sveltos ships an Enterprise tier of its own. Buying the hub is a legitimate scaling decision, the same way managed Kubernetes is.
Who Should Skip This Entire Problem
Not every fleet needs a scaled hub, and pretending otherwise sells you complexity. Skip the sharding conversation entirely if: your fleet is a handful of clusters and hundreds of apps (there is nothing to shard, and tuning at that size mostly adds alpha-featured knobs to operate); your "scaling problem" is actually one giant monorepo (fix the repo structure and the webhook config before touching the controller); or you run a single hub with a spoke fleet of a few small clusters. And if your organization cannot yet operate Redis with HA discipline, do not build a 20,000-app Argo CD hub on it — the fastest way to find out what Redis does is to depend on it at scale.
For choosing between the engines in the first place, read our Argo CD vs. Flux comparison (the reconciliation-loop and drift-trade-off fundamentals apply here too), and for the network topology around a scaled-out hub, our hub-and-spoke vs. mesh cloud networking guide covers the layer below the control plane.
References & Further Reading
- Argo CD — High Availability, Disaster Recovery, and Scaling (the scaling guidance lives inside this page; Argo CD has no standalone scaling doc — verified 404 at time of writing)
- CNOE — Argo CD Benchmarking: Pushing the Limits and Sharding Deep Dive (500 clusters / 50,000 applications; all benchmark numbers in this guide are from this study)
- Argo CD — Dynamic Cluster Distribution (alpha)
- argoproj/argo-cd at v3.5.4 — source of verified defaults: status-processors 20, operation-processors 10, repo-server parallelismlimit 0 (unlimited), Redis 8.2.3
- GHSA-w996-f2wq-x9c6 — oversized OCI manifest can exhaust repo-server memory (fixed in v3.6)
- Sveltos — Sharding (annotation-driven shard sets, shard-components ConfigMap patches)
- Sveltos — Vertical scaling (concurrent-reconciles 10, worker-number 20)
- Sveltos — Architecture (components, modes, no-Redis state model)
- Sveltos v1.16.0 release notes — pull mode with short-lived tokens
- Kubernetes — API Priority and Fairness (the hub's own API-server protection)
- Go — GC guide: memory limits (GOMEMLIMIT)
- Akuity — enterprise Argo CD platform by the Argo CD creators