Long-Term Metrics Storage with Grafana Mimir: Multi-Tenant Prometheus Without the Cardinality Bankruptcy
Sources
- Grafana Mimir documentation
- Grafana Mimir 3.2 release notes (query sharding default-on, remote execution)
- Mimir configuration parameters reference
- Mimir metrics storage retention configuration
- Mimir runtime configuration (per-tenant overrides)
- grafana/mimir GitHub repository (AGPL-3.0)
- Prometheus remote write 2.0 specification
- How we scaled Grafana Mimir to 1 billion active series — Grafana Labs blog
- mimir-distributed Helm chart documentation
Every platform team hits the same wall at the same time. Prometheus is excellent at scraping a cluster and terrible at remembering anything: retention is bounded by the node's disk, the query path lives or dies with one server, and "give us per-team cost and isolation for two hundred clusters" is simply not a question Prometheus was built to answer. The Prometheus design position on this has not moved in a decade — local disk, no clustering, remote write for everything else.
That refusal is why the remote-write ecosystem exists at all, and the two big answers are Thanos and Grafana Mimir. We already covered the VictoriaMetrics alternative in VictoriaMetrics vs Prometheus and the storage-end of the Prometheus 3.x line in Prometheus 3.15.0 release notes. This guide is about Mimir — and specifically about the reason you actually run it: not retention (S3 gives you that), but multi-tenancy with per-tenant cardinality guardrails, enforced in one place, hot-reloaded at runtime.
What we did: ran the real thing. Grafana Mimir 3.2.1 (sha256-verified release binary, go1.26.7) on a single host, plus Prometheus 3.15.0 remote-writing into it as two different tenants. We pushed a 6,000-series burst into a tenant capped at 5,000 series, watched the cap hold exactly, watched the remote-write client drop the excess on the floor, then fixed tenant limits live via hot-reload without touching the server. Every error message, metric value, and config key below is real output from that run — not documentation paraphrase.
1. The Structural Shift: From Local Disk to a Multi-Tenant Facts Store
Prometheus's model is a cache: scrape, keep in TSDB head, compact to local blocks, expire. The ecosystem's remote-write receivers are the system of record. Once metrics leave the cluster, three things you could never get from Prometheus itself become possible — and they are the entire justification for a Mimir deployment:
- Tenancy. One
X-Scope-OrgIDheader separates every write and read. Two Prometheus servers writing to the same Mimir never see each other's series. We verified this live: tenant-a answeredup == 1; the same query under tenant-b returned zero series. - Per-tenant limits as a first-class API. Series caps, ingestion rate, query length, fetched-series-per-query — all set per tenant in a YAML file that reloads every 10 seconds. The noisy-neighbor problem gets a blast radius instead of a pager.
- Cardinality as an observable, budgeted resource. Mimir counts active series per tenant and per custom matcher class, so "who blew up the metrics bill" is a query, not an archaeology dig.
The honest framing: Mimir is the Cortex lineage productized. Grafana forked Mimir from Cortex in March 2022 to pay down accumulated technical debt, kept the battle-tested distributor/ingester/store-gateway split, and made object storage the only source of truth — we traced the history in the Prometheus oss-tale. It scales absurdly (Grafana's own 1-billion-active-series load test, single tenant), it is AGPL-3.0, and it is the engine behind Grafana Cloud's managed metrics. The commercial path exists in two shapes if you don't want to run it: Grafana Cloud (SaaS) and Grafana Enterprise Metrics (self-managed, paid features) — the open source chart deploys either.
Who should skip this guide: if you have one cluster and one team, Mimir is a distributed-systems tax you don't owe. Keep Prometheus local, or use its retention features, and revisit when per-team isolation becomes a real requirement.
2. Architectural Blueprint
Mimir's components map to a single data plane: writes fan out through distributors to a ring of ingesters (in-memory TSDB heads), blocks ship to object storage, and reads go through a query-frontend that splits and caches. The write path and read path have different URL prefixes — a live trap we hit in testing (writes are /api/v1/push, reads are /prometheus/api/v1/...):
Prometheus (cluster A) Prometheus (cluster B)
| remote_write | remote_write
| + header X-Scope-OrgID: team-a | + header X-Scope-OrgID: team-b
v v
+------------- HTTP: /api/v1/push -------------+
| DISTRIBUTOR |
| per-tenant rate limit (429, retryable) |
| replication by consistent-hash ring |
+----------------------+------------------------+
v
+------- INGESTERS -------+ in-memory TSDB head
| ring: N x RF replicas | per-tenant series cap
| (memberlist KV store) | (400, NON-retryable)
+-----------+------------+
| 2h blocks
v
[ OBJECT STORAGE S3 / GCS / Azure ]
^
| compaction + retention
+-----------+------------+ +----------------+
| COMPACTOR | | RULER |
+--------------------------+ | (rule eval, |
| alerting) |
+------------- HTTP: /prometheus/api/v1/ ----+----------------+
| QUERY-FRONTEND | split + cache
| per-tenant query limits, sharding (3.2: on) |
+----------------------+------------------------+
v
QUERIERS ----- store-gateways (bucket index)
|
v
Grafana / your toolsThree architectural facts matter for what follows:
- The ingester, not the distributor, enforces the series cap. That placement is why the cap is exact — the in-memory head is the single place where a tenant's active series are countable.
- Rate limiting (distributor) and series caps (ingester) fail differently — 429 versus 400 — and your remote-write client's response to that difference decides whether you lose samples. Section 4 is the field notes.
- Everything durable lives in object storage. Ingester disks are a cache. There is no "Mimir database node" to lose — there is a bucket, and a compactor that also enforces retention.
3. Complete, Runnable Configuration — Verified on a Live 3.2.1
We ran the full stack on a single host: Mimir 3.2.1 from the official releases page (sha256 8df1ddd5de5a4ad75f3627050b063b19162ba3a20ad98a1e11aafcb0525c988f, matching the published checksum file) and Prometheus 3.15.0. Version output from the binary itself:
$ ./mimir-linux-amd64 --version
Mimir, version 3.2.1 (branch: HEAD, revision: e49585d4)
go version: go1.26.7
platform: linux/amd64The config below is what we actually ran — single-process mode (all components in one binary, -target=all implied), multitenancy left ON because that is the mode that matters, filesystem object storage so the demo needs no bucket. It is derived from the official single-process-config-blocks.yaml shipped in the 3.2.1 tag. Do not run single-process in production; run the Helm chart in section 3.4.
multitenancy_enabled: true
server:
http_listen_port: 9009
grpc_server_max_recv_msg_size: 104857600
grpc_server_max_send_msg_size: 104857600
distributor:
ring:
kvstore:
store: memberlist
pool:
health_check_ingesters: true
ingester:
ring:
min_ready_duration: 0s
final_sleep: 0s
num_tokens: 512
kvstore:
store: inmemory
replication_factor: 1
blocks_storage:
backend: filesystem
filesystem:
dir: /tmp/mimir-demo/blocks
tsdb:
dir: /tmp/mimir-demo/tsdb
bucket_store:
sync_dir: /tmp/mimir-demo/tsdb-sync
compactor:
data_dir: /tmp/mimir-demo/compactor
ruler_storage:
backend: local
local:
directory: /tmp/mimir-demo/rules
runtime_config:
file: /tmp/mimir-demo/overrides.yamlStart it and check the two health endpoints before anything else:
$ ./mimir-linux-amd64 -config.file=mimir.yaml &
# readiness: returns "ready" when the ingester ring is up
$ curl -s http://localhost:9009/ready
ready
# build info: confirms version and feature flags
$ curl -s http://localhost:9009/api/v1/status/buildinfo3.1 Prometheus remote write: the header field everyone gets wrong
Mimir routes every request by tenant header. Prometheus sends that header via the remote-write headers map. The config below is verified working against Prometheus 3.15.0 — and note the field name carefully, because metadata_headers is a remote-write 2.0 spec concept that does not exist as a Prometheus 3.15 config key. Our first attempt used it and got a hard parse error:
err="parsing YAML file prometheus.yml: yaml: unmarshal errors:
line 13: field metadata_headers not found in type config.plain"The correct key is headers (defined on RemoteWriteConfig in Prometheus's config.go, validated against a reserved list — X-Scope-OrgID is allowed):
global:
scrape_interval: 1s
evaluation_interval: 1s
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ['localhost:9090']
remote_write:
- url: http://localhost:9009/api/v1/push
headers:
X-Scope-OrgID: tenant-a
queue_config:
capacity: 10000
max_shards: 5Within seconds, Mimir's log showed ingestion under the tenant identity — the confirmation that the pipeline is wired correctly:
ts=2026-10-05T05:09:38.326Z caller=handler.go:95 level=info user=tenant-a
caller=head.go:966 time=2026-10-05T05:09:38.326Z msg="WAL segment loaded" segment=0
ts=2026-10-05T05:09:38.326Z caller=ingester_tsdb.go:252 level=info user=tenant-a
msg="Running compaction after WAL replay"3.2 Querying: the /prometheus prefix and the missing-header error
Mimir exposes the Prometheus read API under /prometheus, not at the root. Queries to /api/v1/query return 404 page not found — our own first probe did exactly that. The write endpoint is /api/v1/push; reads are /prometheus/api/v1/.... And without the tenant header, the API refuses with a bare no org id — that error text is Mimir telling you multitenancy is on and your client forgot who it is:
# correct read path, tenant-a: the scraped up metric is there
$ curl -s -H "X-Scope-OrgID: tenant-a" \
"http://localhost:9009/prometheus/api/v1/query?query=up"
{"status":"success","data":{"resultType":"vector","result":[
{"metric":{"instance":"localhost:9090","job":"prometheus"},
"value":[1791177032.09,"1"]}]}}
# tenant-b: same query, zero series — isolation holds
$ curl -s -H "X-Scope-OrgID: tenant-b" \
"http://localhost:9009/prometheus/api/v1/query?query=up"
{"status":"success","data":{"resultType":"vector","result":[]}}
# no header: rejected outright
$ curl -s "http://localhost:9009/prometheus/api/v1/query?query=up"
no org id3.3 The cardinality demo: a 6,000-series burst into a 5,000-series tenant
This is the part most Mimir articles skip because it requires breaking things. We stood up a metrics exporter serving 6,000 unique series (a cardinality_probe{series_id="0".."5999"} family) and pointed a second Prometheus at it, remote-writing to tenant-b, whose override capped max_global_series_per_user at 5,000. The overrides file — and note the exact schema, because the schema is where we initially failed (see 3.5):
overrides:
tenant-a:
ingestion_rate: 5000
max_global_series_per_user: 100000
tenant-b:
ingestion_rate: 1000
max_global_series_per_user: 5000
cardinality_analysis_enabled: true
active_series_custom_trackers:
probe-series: '{__name__="cardinality_probe"}'The result, from Mimir's own metrics after the burst: the cap held at exactly 5,000 active series, and everything above it was counted as discarded:
$ curl -s http://localhost:9009/metrics | grep -E "^cortex_"
cortex_ingester_active_series{user="tenant-a"} 761
cortex_ingester_active_series{user="tenant-b"} 5000
cortex_ingester_active_series_custom_tracker{name="probe-series",user="tenant-b"} 5000
cortex_discarded_samples_total{reason="per_user_series_limit",user="tenant-b"} 214560
cortex_discarded_samples_total{reason="rate_limited",user="tenant-b"} 900000And the rejection itself, exactly as Mimir logged it (sampled 1-in-10):
msg="detected an error while ingesting Prometheus remote-write request
(the request may have been partially ingested)"
err="send data to ingesters: failed pushing to ingester pm: user=tenant-b:
per-user series limit of 5000 exceeded (err-mimir-max-series-per-user).
To adjust the related per-tenant limit, configure
-ingester.max-global-series-per-user, or contact your service administrator."Yes, Mimir still exports cortex_* metric names — the fork kept the lineage, dashboards included. The active_series_custom_trackers key is the modern per-tenant tracker: a flat map of tracker-name: matcher (Grafana's own example is dev: '{namespace=~"dev-.*"}'), which shows up as its own metric — that custom_tracker{name="probe-series"} line above is the tracker we configured live.
3.4 Production: the Helm chart, not the single binary
For real deployments, Mimir ships as the mimir-distributed chart (version 6.3.0-weekly.413 as of this writing, requires Kubernetes ≥ 1.32, deploys Mimir or Grafana Enterprise Metrics). The chart moved to operations/helm/charts/mimir-distributed in the repo — old blog posts pointing at operations/helm-charts/ are stale:
kubectl create namespace mimir-test
helm repo add grafana https://grafana.github.io/helm-charts
helm repo update
helm -n mimir-test install mimir grafana/mimir-distributedThree things the chart does that you should not undo: zone-aware replication is on by default for new installs (ingesters spread across zones, replication_factor: 3); MinIO is disabled by default — you are expected to bring S3/GCS/Azure; and the chart's own size presets are honest engineering references (the capped-small.yaml profile is sized for ~1M series at a 15s scrape interval ≈ 66,000 samples/s, with requests==limits for Guaranteed QoS). Wire the chart to your bucket via mimir.structuredConfig.blocks_storage.s3 — the object storage configuration docs cover credentials, SSE, and endpoint shapes. For per-tenant limits in the chart, keep the same overrides structure but manage the file where your GitOps pipeline can review it — runtime YAML is still YAML; treat it as code.
4. Day-2 Operational Warnings
Everything in this section cost us real minutes during the run. It is the part of the guide to read twice.
4.1 The 400 trap: series caps are non-retryable, and your client silently drops data
The single most important thing to understand about Mimir limits is that they fail with different HTTP codes, and Prometheus remote write treats them oppositely. When tenant-b hit its series cap, Mimir returned HTTP 400. Prometheus classifies 400 as non-recoverable — the samples are dropped from the queue permanently:
level=ERROR source=queue_manager.go:1731 msg="non-recoverable error"
component=remote remote_name=28268f url=http://localhost:9009/api/v1/push
failedSampleCount=2000 failedHistogramCount=0 failedExemplarCount=0
err="server returned HTTP status 400 Bad Request: send data to ingesters:
failed pushing to ingester pm: user=tenant-b: per-user series limit of 5000
exceeded (err-mimir-max-series-per-user) ..."Contrast the ingestion rate limit, which Mimir returns as HTTP 429 — Prometheus backs off and retries, and the samples eventually land. We hit that one too (tenant-b was capped at 1,000 items/s and we pushed well past it):
msg="detected an error while ingesting Prometheus remote-write request
(the request may have been partially ingested)" httpCode=429
err="the request has been rejected because the tenant exceeded the ingestion
rate limit, set to 1000 items/s with a maximum allowed burst of 200000.
This limit is applied on the total number of samples, exemplars and metadata
received across all distributors (err-mimir-tenant-max-ingestion-rate).
To adjust the related per-tenant limits, configure
-distributor.ingestion-rate-limit and -distributor.ingestion-burst-size,
or contact your service administrator."| Limit | Enforced in | HTTP code | Remote-write client behavior | Production consequence |
|---|---|---|---|---|
ingestion_rate (default 10,000/s) | Distributor | 429 | Retry with backoff | Latency spike, no loss |
max_global_series_per_user (default 150,000) | Ingester | 400 | Drop permanently | Silent data loss for new series |
max_active_series_per_user (newer, 0 = off) | Ingester | 429 (configurable via active_series_limit_response_code) | Retry | Lossless backpressure |
max_fetched_series_per_query | Query-frontend | 422 | Query fails | Dashboard error, no ingestion impact |
That third row is your migration path: the newer max_active_series_per_user limit exists precisely because a 400 on series caps is hostile, and it defaults to a retryable 429 (active_series_limit_response_code: 429 in the effective-tenant dump below). But the old global cap is still the one most fleets have on. Monitor the discard counter and alert on it — this is your data-loss canary:
# any nonzero value here means a tenant is LOSING samples right now
sum by (user, reason) (rate(cortex_discarded_samples_total[5m]))
# the ten tenants burning the most cardinality right now —
# compare against each tenant's configured max_global_series_per_user
topk(10, cortex_ingester_active_series)4.2 Hot reload semantics: 10-second pickup, last-good-config on error
Runtime overrides reload every -runtime-config.reload-period — default 10 seconds — with no restart, no SIGHUP, no API call. We verified both directions:
- Good file: flipping
cardinality_analysis_enabled: trueand adding a custom tracker for tenant-b took effect within one reload cycle — the effective-tenant config endpoint (GET /runtime_config) confirmed it, and the tracker's metric appeared immediately. - Broken file: we fat-fingered the tracker schema (nested
matcher:key under the tracker name — it must be a flatname: matchermap). Mimir logged"failed to load config"every 10 seconds, 44 times, and kept serving the last good config the entire time. No crash, no rollback notification — just an error loop you must be alerting on.
ts=2026-10-05T05:15:01Z caller=manager.go:193 level=error
msg="failed to load config" err="load file: yaml: unmarshal errors:
line 8: cannot unmarshal !!map into string"
ts=2026-10-05T05:15:11Z caller=manager.go:193 level=error
msg="failed to load config" ...Operational rule: alert on the log pattern failed to load config (the reload manager logs it per cycle — it never exits), and treat GET /runtime_config as your drift detector between what Git says and what Mimir is actually enforcing. The endpoint dumps every effective per-tenant limit — we used it to catch our own broken override in minutes.
4.3 The cardinality analysis endpoints are off by default
/prometheus/api/v1/cardinality/label_values and friends 404 until you set cardinality_analysis_enabled: true per tenant — it is a query-path cost control, not a global default. Once on, the label-values endpoint demands label_names[] (array form) and the active-series endpoint demands selector — both validated live:
$ curl -s -H "X-Scope-OrgID: tenant-b" \
"http://localhost:9009/prometheus/api/v1/cardinality/label_values?label_names[]=series_id&start=1791174000&end=1791177000"
{"series_count_total":5000,"labels":[
{"label_name":"series_id","label_values_count":5000,"series_count":5000,
"cardinality":[{"label_value":"0","series_count":1},
{"label_value":"1","series_count":1}, ...]}]}
# wrong params are rejected helpfully:
# "'label_names[]' param is required"
# "selector parameter is required"4.4 Upgrade and config landmines in 3.2
- Query sharding is now default-on (
-query-frontend.parallelize-shardable-queries). Existing query load will fan out differently after upgrade — watch querier CPU and results-cache hit rates. - Remote execution is default-on and version-gated: all queriers must be on 3.1 before you upgrade the frontend to 3.2. The old compatibility flag is gone. Read the 3.2 release notes before touching the chart.
- Config key rename:
max_query_lengthis nowmax_total_query_length. Our config with the old key failed to parse —"field max_query_length not found in type validation.plainLimits". If your generated configs still emit the old name, they will not apply. - Ingester-to-querier request hedging is now default-off (was 3s). Tail latencies may shift; the old behavior returns with
-querier.minimize-ingester-requests-hedging-delay=3s. - Compactor split-and-merge shards are rounded up to a power of two — an existing odd value silently changes compaction topology.
- Defaults worth knowing cold: 150,000 series/tenant, 10,000 samples/s ingestion with 200,000 burst, 24h query split interval, 30 label names per series, 10-minute creation grace for "new" series detection. All verified in the live
/runtime_configdump against the configuration parameters reference.
4.5 Metrics to watch, blast radius notes
| Metric | What it tells you | Alarm posture |
|---|---|---|
cortex_discarded_samples_total{reason="per_user_series_limit"} | Active data loss from series caps | Page — this is the 400 trap firing |
cortex_discarded_samples_total{reason="rate_limited"} | Tenant over rate budget (429, retried) | Warn + raise tenant's budget review |
cortex_ingester_active_series{user=...} | Per-tenant cardinality burn rate | Capacity planning + tenant chargeback |
cortex_ingester_active_series_custom_tracker{name=...} | Cardinality of a matcher class you care about (e.g. a noisy app) | Tenant-specific SLOs |
log pattern "failed to load config" (manager.go) | Runtime override file is invalid | Warn — Mimir is frozen on last-good |
Blast radius: a bad overrides file cannot corrupt state (last-good config keeps serving), but a bad tenant can — 150,000 default series per tenant across N tenants is the compounding risk the caps exist for. Set explicit per-tenant caps before onboarding the second tenant, not after the first incident. The Tenants and Top Tenants dashboards in the official mixin are the fastest way to see which tenant is the problem.
If you run the ruler: rules are per-tenant too (ruler_max_rules_per_rule_group: 20, ruler_max_rule_groups_per_tenant: 70 defaults), and the ruler_evaluation_delay_duration default of 1m protects read-your-writes at the cost of alert latency. If you pipe metrics through an OpenTelemetry Collector before Mimir, the OTel ingestion path has its own label-translation flags — review them before assuming parity with the Prometheus wire format.
5. Who Should Run Mimir, Who Should Not
Run Mimir when: you operate metrics for multiple teams/clusters and need hard isolation; per-tenant cost attribution and cardinality budgets are real requirements; you already run S3/GCS and a Kubernetes fleet; you want the option of Grafana Cloud as a burst valve (same engine, same API) or GEM for paid features with self-hosting. Our Grafana Cloud vs Datadog pricing guide covers when the managed version of this stack wins.
Skip Mimir when: one cluster, one team — local Prometheus or VictoriaMetrics is less operational surface; your retention need is modest — Thanos sidecar + S3 does retention without a microservices fleet; you cannot commit to operating object storage properly — Mimir without bucket versioning/Lifecycle rules is a data-durability liability. And if none of your tenants will ever exceed ~150k series, the default caps will make you feel limits you never needed — the fleet is the point.
References & Further Reading
- Grafana Mimir documentation — architecture, configuration, and operations reference
- Mimir configuration parameters — every flag and default, including the 3.2 renames
- Runtime configuration — per-tenant overrides and reload semantics
- Metrics storage retention — compactor-enforced retention configuration
- mimir-distributed Helm chart — production deployment reference
- grafana/mimir on GitHub — source, CHANGELOG, and release checksums (AGPL-3.0)
- Prometheus remote write 2.0 specification — wire protocol Mimir ingests
- How we scaled Grafana Mimir to 1 billion active series — Grafana Labs load-test write-up
- Grafana Cloud managed metrics — the commercial path running the same engine
- Platform Monkey: The Prometheus oss-tale · Prometheus 3.15.0 release notes · VictoriaMetrics vs Prometheus · OTel Collector architecture · Grafana Cloud vs Datadog