Long-Term Metrics Storage with Grafana Mimir: Multi-Tenant Prometheus Without the Cardinality Bankruptcy

Sources

Every platform team hits the same wall at the same time. Prometheus is excellent at scraping a cluster and terrible at remembering anything: retention is bounded by the node's disk, the query path lives or dies with one server, and "give us per-team cost and isolation for two hundred clusters" is simply not a question Prometheus was built to answer. The Prometheus design position on this has not moved in a decade — local disk, no clustering, remote write for everything else.

That refusal is why the remote-write ecosystem exists at all, and the two big answers are Thanos and Grafana Mimir. We already covered the VictoriaMetrics alternative in VictoriaMetrics vs Prometheus and the storage-end of the Prometheus 3.x line in Prometheus 3.15.0 release notes. This guide is about Mimir — and specifically about the reason you actually run it: not retention (S3 gives you that), but multi-tenancy with per-tenant cardinality guardrails, enforced in one place, hot-reloaded at runtime.

What we did: ran the real thing. Grafana Mimir 3.2.1 (sha256-verified release binary, go1.26.7) on a single host, plus Prometheus 3.15.0 remote-writing into it as two different tenants. We pushed a 6,000-series burst into a tenant capped at 5,000 series, watched the cap hold exactly, watched the remote-write client drop the excess on the floor, then fixed tenant limits live via hot-reload without touching the server. Every error message, metric value, and config key below is real output from that run — not documentation paraphrase.

1. The Structural Shift: From Local Disk to a Multi-Tenant Facts Store

Prometheus's model is a cache: scrape, keep in TSDB head, compact to local blocks, expire. The ecosystem's remote-write receivers are the system of record. Once metrics leave the cluster, three things you could never get from Prometheus itself become possible — and they are the entire justification for a Mimir deployment:

The honest framing: Mimir is the Cortex lineage productized. Grafana forked Mimir from Cortex in March 2022 to pay down accumulated technical debt, kept the battle-tested distributor/ingester/store-gateway split, and made object storage the only source of truth — we traced the history in the Prometheus oss-tale. It scales absurdly (Grafana's own 1-billion-active-series load test, single tenant), it is AGPL-3.0, and it is the engine behind Grafana Cloud's managed metrics. The commercial path exists in two shapes if you don't want to run it: Grafana Cloud (SaaS) and Grafana Enterprise Metrics (self-managed, paid features) — the open source chart deploys either.

Who should skip this guide: if you have one cluster and one team, Mimir is a distributed-systems tax you don't owe. Keep Prometheus local, or use its retention features, and revisit when per-team isolation becomes a real requirement.

2. Architectural Blueprint

Mimir's components map to a single data plane: writes fan out through distributors to a ring of ingesters (in-memory TSDB heads), blocks ship to object storage, and reads go through a query-frontend that splits and caches. The write path and read path have different URL prefixes — a live trap we hit in testing (writes are /api/v1/push, reads are /prometheus/api/v1/...):

   Prometheus (cluster A)          Prometheus (cluster B)
        | remote_write                     | remote_write
        | + header X-Scope-OrgID: team-a  | + header X-Scope-OrgID: team-b
        v                                 v
   +------------- HTTP: /api/v1/push -------------+
   |                  DISTRIBUTOR                  |
   |  per-tenant rate limit (429, retryable)       |
   |  replication by consistent-hash ring          |
   +----------------------+------------------------+
                          v
                  +------- INGESTERS -------+   in-memory TSDB head
                  | ring: N x RF replicas   |   per-tenant series cap
                  | (memberlist KV store)   |   (400, NON-retryable)
                  +-----------+------------+
                              | 2h blocks
                              v
                  [ OBJECT STORAGE  S3 / GCS / Azure ]
                              ^
                              | compaction + retention
                  +-----------+------------+   +----------------+
                  |       COMPACTOR         |   |    RULER       |
                  +--------------------------+   | (rule eval,    |
                                                 |  alerting)     |
   +------------- HTTP: /prometheus/api/v1/ ----+----------------+
   |                QUERY-FRONTEND                 |  split + cache
   |  per-tenant query limits, sharding (3.2: on)  |
   +----------------------+------------------------+
                          v
                     QUERIERS  ----- store-gateways (bucket index)
                          |
                          v
                    Grafana / your tools

Three architectural facts matter for what follows:

3. Complete, Runnable Configuration — Verified on a Live 3.2.1

We ran the full stack on a single host: Mimir 3.2.1 from the official releases page (sha256 8df1ddd5de5a4ad75f3627050b063b19162ba3a20ad98a1e11aafcb0525c988f, matching the published checksum file) and Prometheus 3.15.0. Version output from the binary itself:

$ ./mimir-linux-amd64 --version
Mimir, version 3.2.1 (branch: HEAD, revision: e49585d4)
  go version:       go1.26.7
  platform:         linux/amd64

The config below is what we actually ran — single-process mode (all components in one binary, -target=all implied), multitenancy left ON because that is the mode that matters, filesystem object storage so the demo needs no bucket. It is derived from the official single-process-config-blocks.yaml shipped in the 3.2.1 tag. Do not run single-process in production; run the Helm chart in section 3.4.

multitenancy_enabled: true

server:
  http_listen_port: 9009
  grpc_server_max_recv_msg_size: 104857600
  grpc_server_max_send_msg_size: 104857600

distributor:
  ring:
    kvstore:
      store: memberlist
  pool:
    health_check_ingesters: true

ingester:
  ring:
    min_ready_duration: 0s
    final_sleep: 0s
    num_tokens: 512
    kvstore:
      store: inmemory
    replication_factor: 1

blocks_storage:
  backend: filesystem
  filesystem:
    dir: /tmp/mimir-demo/blocks
  tsdb:
    dir: /tmp/mimir-demo/tsdb
  bucket_store:
    sync_dir: /tmp/mimir-demo/tsdb-sync

compactor:
  data_dir: /tmp/mimir-demo/compactor

ruler_storage:
  backend: local
  local:
    directory: /tmp/mimir-demo/rules

runtime_config:
  file: /tmp/mimir-demo/overrides.yaml

Start it and check the two health endpoints before anything else:

$ ./mimir-linux-amd64 -config.file=mimir.yaml &

# readiness: returns "ready" when the ingester ring is up
$ curl -s http://localhost:9009/ready
ready

# build info: confirms version and feature flags
$ curl -s http://localhost:9009/api/v1/status/buildinfo

3.1 Prometheus remote write: the header field everyone gets wrong

Mimir routes every request by tenant header. Prometheus sends that header via the remote-write headers map. The config below is verified working against Prometheus 3.15.0 — and note the field name carefully, because metadata_headers is a remote-write 2.0 spec concept that does not exist as a Prometheus 3.15 config key. Our first attempt used it and got a hard parse error:

err="parsing YAML file prometheus.yml: yaml: unmarshal errors:
  line 13: field metadata_headers not found in type config.plain"

The correct key is headers (defined on RemoteWriteConfig in Prometheus's config.go, validated against a reserved list — X-Scope-OrgID is allowed):

global:
  scrape_interval: 1s
  evaluation_interval: 1s
scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets: ['localhost:9090']
remote_write:
  - url: http://localhost:9009/api/v1/push
    headers:
      X-Scope-OrgID: tenant-a
    queue_config:
      capacity: 10000
      max_shards: 5

Within seconds, Mimir's log showed ingestion under the tenant identity — the confirmation that the pipeline is wired correctly:

ts=2026-10-05T05:09:38.326Z caller=handler.go:95 level=info user=tenant-a
  caller=head.go:966 time=2026-10-05T05:09:38.326Z msg="WAL segment loaded" segment=0
ts=2026-10-05T05:09:38.326Z caller=ingester_tsdb.go:252 level=info user=tenant-a
  msg="Running compaction after WAL replay"

3.2 Querying: the /prometheus prefix and the missing-header error

Mimir exposes the Prometheus read API under /prometheus, not at the root. Queries to /api/v1/query return 404 page not found — our own first probe did exactly that. The write endpoint is /api/v1/push; reads are /prometheus/api/v1/.... And without the tenant header, the API refuses with a bare no org id — that error text is Mimir telling you multitenancy is on and your client forgot who it is:

# correct read path, tenant-a: the scraped up metric is there
$ curl -s -H "X-Scope-OrgID: tenant-a" \
    "http://localhost:9009/prometheus/api/v1/query?query=up"
{"status":"success","data":{"resultType":"vector","result":[
  {"metric":{"instance":"localhost:9090","job":"prometheus"},
   "value":[1791177032.09,"1"]}]}}

# tenant-b: same query, zero series — isolation holds
$ curl -s -H "X-Scope-OrgID: tenant-b" \
    "http://localhost:9009/prometheus/api/v1/query?query=up"
{"status":"success","data":{"resultType":"vector","result":[]}}

# no header: rejected outright
$ curl -s "http://localhost:9009/prometheus/api/v1/query?query=up"
no org id

3.3 The cardinality demo: a 6,000-series burst into a 5,000-series tenant

This is the part most Mimir articles skip because it requires breaking things. We stood up a metrics exporter serving 6,000 unique series (a cardinality_probe{series_id="0".."5999"} family) and pointed a second Prometheus at it, remote-writing to tenant-b, whose override capped max_global_series_per_user at 5,000. The overrides file — and note the exact schema, because the schema is where we initially failed (see 3.5):

overrides:
  tenant-a:
    ingestion_rate: 5000
    max_global_series_per_user: 100000
  tenant-b:
    ingestion_rate: 1000
    max_global_series_per_user: 5000
    cardinality_analysis_enabled: true
    active_series_custom_trackers:
      probe-series: '{__name__="cardinality_probe"}'

The result, from Mimir's own metrics after the burst: the cap held at exactly 5,000 active series, and everything above it was counted as discarded:

$ curl -s http://localhost:9009/metrics | grep -E "^cortex_"

cortex_ingester_active_series{user="tenant-a"} 761
cortex_ingester_active_series{user="tenant-b"} 5000
cortex_ingester_active_series_custom_tracker{name="probe-series",user="tenant-b"} 5000
cortex_discarded_samples_total{reason="per_user_series_limit",user="tenant-b"} 214560
cortex_discarded_samples_total{reason="rate_limited",user="tenant-b"} 900000

And the rejection itself, exactly as Mimir logged it (sampled 1-in-10):

msg="detected an error while ingesting Prometheus remote-write request
(the request may have been partially ingested)"
err="send data to ingesters: failed pushing to ingester pm: user=tenant-b:
per-user series limit of 5000 exceeded (err-mimir-max-series-per-user).
To adjust the related per-tenant limit, configure
-ingester.max-global-series-per-user, or contact your service administrator."

Yes, Mimir still exports cortex_* metric names — the fork kept the lineage, dashboards included. The active_series_custom_trackers key is the modern per-tenant tracker: a flat map of tracker-name: matcher (Grafana's own example is dev: '{namespace=~"dev-.*"}'), which shows up as its own metric — that custom_tracker{name="probe-series"} line above is the tracker we configured live.

3.4 Production: the Helm chart, not the single binary

For real deployments, Mimir ships as the mimir-distributed chart (version 6.3.0-weekly.413 as of this writing, requires Kubernetes ≥ 1.32, deploys Mimir or Grafana Enterprise Metrics). The chart moved to operations/helm/charts/mimir-distributed in the repo — old blog posts pointing at operations/helm-charts/ are stale:

kubectl create namespace mimir-test
helm repo add grafana https://grafana.github.io/helm-charts
helm repo update
helm -n mimir-test install mimir grafana/mimir-distributed

Three things the chart does that you should not undo: zone-aware replication is on by default for new installs (ingesters spread across zones, replication_factor: 3); MinIO is disabled by default — you are expected to bring S3/GCS/Azure; and the chart's own size presets are honest engineering references (the capped-small.yaml profile is sized for ~1M series at a 15s scrape interval ≈ 66,000 samples/s, with requests==limits for Guaranteed QoS). Wire the chart to your bucket via mimir.structuredConfig.blocks_storage.s3 — the object storage configuration docs cover credentials, SSE, and endpoint shapes. For per-tenant limits in the chart, keep the same overrides structure but manage the file where your GitOps pipeline can review it — runtime YAML is still YAML; treat it as code.

4. Day-2 Operational Warnings

Everything in this section cost us real minutes during the run. It is the part of the guide to read twice.

4.1 The 400 trap: series caps are non-retryable, and your client silently drops data

The single most important thing to understand about Mimir limits is that they fail with different HTTP codes, and Prometheus remote write treats them oppositely. When tenant-b hit its series cap, Mimir returned HTTP 400. Prometheus classifies 400 as non-recoverable — the samples are dropped from the queue permanently:

level=ERROR source=queue_manager.go:1731 msg="non-recoverable error"
component=remote remote_name=28268f url=http://localhost:9009/api/v1/push
failedSampleCount=2000 failedHistogramCount=0 failedExemplarCount=0
err="server returned HTTP status 400 Bad Request: send data to ingesters:
failed pushing to ingester pm: user=tenant-b: per-user series limit of 5000
exceeded (err-mimir-max-series-per-user) ..."

Contrast the ingestion rate limit, which Mimir returns as HTTP 429 — Prometheus backs off and retries, and the samples eventually land. We hit that one too (tenant-b was capped at 1,000 items/s and we pushed well past it):

msg="detected an error while ingesting Prometheus remote-write request
(the request may have been partially ingested)" httpCode=429
err="the request has been rejected because the tenant exceeded the ingestion
rate limit, set to 1000 items/s with a maximum allowed burst of 200000.
This limit is applied on the total number of samples, exemplars and metadata
received across all distributors (err-mimir-tenant-max-ingestion-rate).
To adjust the related per-tenant limits, configure
-distributor.ingestion-rate-limit and -distributor.ingestion-burst-size,
or contact your service administrator."
LimitEnforced inHTTP codeRemote-write client behaviorProduction consequence
ingestion_rate (default 10,000/s)Distributor429Retry with backoffLatency spike, no loss
max_global_series_per_user (default 150,000)Ingester400Drop permanentlySilent data loss for new series
max_active_series_per_user (newer, 0 = off)Ingester429 (configurable via active_series_limit_response_code)RetryLossless backpressure
max_fetched_series_per_queryQuery-frontend422Query failsDashboard error, no ingestion impact

That third row is your migration path: the newer max_active_series_per_user limit exists precisely because a 400 on series caps is hostile, and it defaults to a retryable 429 (active_series_limit_response_code: 429 in the effective-tenant dump below). But the old global cap is still the one most fleets have on. Monitor the discard counter and alert on it — this is your data-loss canary:

# any nonzero value here means a tenant is LOSING samples right now
sum by (user, reason) (rate(cortex_discarded_samples_total[5m]))

# the ten tenants burning the most cardinality right now —
# compare against each tenant's configured max_global_series_per_user
topk(10, cortex_ingester_active_series)

4.2 Hot reload semantics: 10-second pickup, last-good-config on error

Runtime overrides reload every -runtime-config.reload-period — default 10 seconds — with no restart, no SIGHUP, no API call. We verified both directions:

ts=2026-10-05T05:15:01Z caller=manager.go:193 level=error
  msg="failed to load config" err="load file: yaml: unmarshal errors:
  line 8: cannot unmarshal !!map into string"
ts=2026-10-05T05:15:11Z caller=manager.go:193 level=error
  msg="failed to load config" ...

Operational rule: alert on the log pattern failed to load config (the reload manager logs it per cycle — it never exits), and treat GET /runtime_config as your drift detector between what Git says and what Mimir is actually enforcing. The endpoint dumps every effective per-tenant limit — we used it to catch our own broken override in minutes.

4.3 The cardinality analysis endpoints are off by default

/prometheus/api/v1/cardinality/label_values and friends 404 until you set cardinality_analysis_enabled: true per tenant — it is a query-path cost control, not a global default. Once on, the label-values endpoint demands label_names[] (array form) and the active-series endpoint demands selector — both validated live:

$ curl -s -H "X-Scope-OrgID: tenant-b" \
  "http://localhost:9009/prometheus/api/v1/cardinality/label_values?label_names[]=series_id&start=1791174000&end=1791177000"
{"series_count_total":5000,"labels":[
  {"label_name":"series_id","label_values_count":5000,"series_count":5000,
   "cardinality":[{"label_value":"0","series_count":1},
                  {"label_value":"1","series_count":1}, ...]}]}

# wrong params are rejected helpfully:
# "'label_names[]' param is required"
# "selector parameter is required"

4.4 Upgrade and config landmines in 3.2

4.5 Metrics to watch, blast radius notes

MetricWhat it tells youAlarm posture
cortex_discarded_samples_total{reason="per_user_series_limit"}Active data loss from series capsPage — this is the 400 trap firing
cortex_discarded_samples_total{reason="rate_limited"}Tenant over rate budget (429, retried)Warn + raise tenant's budget review
cortex_ingester_active_series{user=...}Per-tenant cardinality burn rateCapacity planning + tenant chargeback
cortex_ingester_active_series_custom_tracker{name=...}Cardinality of a matcher class you care about (e.g. a noisy app)Tenant-specific SLOs
log pattern "failed to load config" (manager.go)Runtime override file is invalidWarn — Mimir is frozen on last-good

Blast radius: a bad overrides file cannot corrupt state (last-good config keeps serving), but a bad tenant can — 150,000 default series per tenant across N tenants is the compounding risk the caps exist for. Set explicit per-tenant caps before onboarding the second tenant, not after the first incident. The Tenants and Top Tenants dashboards in the official mixin are the fastest way to see which tenant is the problem.

If you run the ruler: rules are per-tenant too (ruler_max_rules_per_rule_group: 20, ruler_max_rule_groups_per_tenant: 70 defaults), and the ruler_evaluation_delay_duration default of 1m protects read-your-writes at the cost of alert latency. If you pipe metrics through an OpenTelemetry Collector before Mimir, the OTel ingestion path has its own label-translation flags — review them before assuming parity with the Prometheus wire format.

5. Who Should Run Mimir, Who Should Not

Run Mimir when: you operate metrics for multiple teams/clusters and need hard isolation; per-tenant cost attribution and cardinality budgets are real requirements; you already run S3/GCS and a Kubernetes fleet; you want the option of Grafana Cloud as a burst valve (same engine, same API) or GEM for paid features with self-hosting. Our Grafana Cloud vs Datadog pricing guide covers when the managed version of this stack wins.

Skip Mimir when: one cluster, one team — local Prometheus or VictoriaMetrics is less operational surface; your retention need is modest — Thanos sidecar + S3 does retention without a microservices fleet; you cannot commit to operating object storage properly — Mimir without bucket versioning/Lifecycle rules is a data-durability liability. And if none of your tenants will ever exceed ~150k series, the default caps will make you feel limits you never needed — the fleet is the point.

References & Further Reading