Prometheus: The Monitoring Monoculture That Refused to Cluster
Ex-Googlers at SoundCloud, missing Borgmon, built the pull-based monitoring monoculture that graduated second from the CNCF — and stayed single-node on purpose.
Sources
- prometheus/prometheus repository (created 2012-11-24)
- Prometheus: Monitoring at SoundCloud — Julius Volz & Björn Rabenstein (2015-01-26)
- Initial commit 734d28b515 — 'Initial commit', Matt T. Proud (2012-11-24)
- CNCF accepts Prometheus as second hosted project (2016-05-09)
- Prometheus to join the CNCF — project blog (2016-05-09)
- CNCF announces Prometheus graduation (2018-08-09, PromCon)
- Wikipedia: Prometheus (software) — history and version section
- LWN: Changes in Prometheus 2.0 — TSDB rewrite numbers
- Tom Wilkie, Prometheus 2.0 (GrafanaCon 2018) — TSDB v3 slides
- Prometheus 3.0 release notes (2024-11-14)
- Prometheus 3.15.0 release notes (2026-09-25) — OM2.0 scrape support
- PR #18606: implement OM2 scrape format (merged 2026-08-27)
- Issue #12024: Support Unix socket for metrics address (2023-02-27, closed)
- CNCF TOC issue #1364: [ARCHIVE] (Migrate) OpenMetrics
- OpenMetrics 2.0 WG announcement (prometheus-developers list)
- OpenMetrics repository moved to github.com/prometheus (2024-10-14)
- Cortex+Thanos origin post — Grafana blog (PromCon 2019 recap)
- Announcing Grafana Mimir (2022-03)
- prometheus/alertmanager repository (created 2013-07-16)
- prometheus.io — official documentation root
- thanos.io — CNCF incubating long-term-storage project
- Prometheus exposition formats documentation
- Remote Write 2.0 specification
- SRE Book chapter: Practical Alerting (Borgmon lineage)
- Prometheus 3.0 migration guide
- v2.40.0 release notes — experimental native histograms (#11447)
- CNCF project page — Prometheus
Every ecosystem has one tool whose name became the category. For cloud-native metrics, that tool is Prometheus — and its story is the story of an entire industry learning to monitor what containers broke. It starts with two SoundCloud engineers missing an internal Google system they could no longer use, and ends — for now — with a fourteen-year-old codebase shipping its own wire format into the spec that was once spun out to escape it.
The Origin: Borgmon Withdrawal at SoundCloud
Ex-Googlers, StatsD pain, and a 2012 pet project
The design DNA of Prometheus is not original — it is Borgmon, Google's internal time-series monitoring system, described publicly in the SRE Book's chapter on practical alerting. Matt T. Proud and Julius Volz had both worked at Google and left. At SoundCloud, the pair hit the wall every fast-growing microservice shop hit in 2012: hundreds of services, thousands of instances churning constantly, and a StatsD+Graphite stack that, in their own words, ran into "a number of serious limitations." The founding SoundCloud blog post — written by Julius Volz and Björn Rabenstein in January 2015 — describes the wish list exactly: multi-dimensional data model, operational simplicity, decentralized collection, and a powerful query language, all in one system. Their conclusion: "we could not identify a system that combined them all until a colleague started an ambitious pet project in 2012 that aimed to do so."
The git archaeology confirms the story. The repository was created on , and the deepest commit page reads:
# Deepest page of prometheus/prometheus commit history:
$ gh api "repos/prometheus/prometheus/commits?until=2012-12-31T23:59:59Z&per_page=100"
DATE SHA AUTHOR MESSAGE
2012-11-24 734d28b515 Matt T. Proud Initial commit
2012-12-09 577acf4fe7 Matt T. Proud Exploding the storage infrastructure by contexts
2012-12-11 6589fc92f8 Matt T. Proud Strip web services, which weren't adding value
2012-12-12 59a708f25a Matt T. Proud Provide prototype of storage layer interfaces
2012-12-12 0886592ebc Matt T. Proud New interface definition after discussion
2012-12-19 a14dbd5bd0 Matt T. Proud Interim commit for Julius
# Every commit on that first page: Matt T. Proud, prototyping interfaces,
# and handing interim state to Julius Volz. A pet project, in public.
SoundCloud put it into production by 2013 — before any announcement — and the alertmanager
repository followed on , split out early because alerting is a
different lifecycle problem from collection. The project went public in January 2015 with the SoundCloud blog
post above, the first PromCon followed, and the rest is landscape history: in May 2016 the CNCF's TOC voted
unanimously to accept Prometheus as the foundation's second hosted project after Kubernetes —
and in August 2018, at PromCon, it became the second project to graduate, after Kubernetes.
Not bad for a monitoring tool. Not bad at all.
The Timeline: From Pet Project to Graduated Standard
- 2012-11-24 —
prometheus/prometheuscreated. First commit734d28b515by Matt T. Proud ("Initial commit"). A SoundCloud pet project becomes the company's monitoring system. - 2013 — Production use at SoundCloud (per the founding team's own account and Wikipedia's history). Design influences: Borgmon's pull model, multi-dimensional labels, and the insight that instrumentation belongs in your own code.
- 2013-07-16 —
prometheus/alertmanagerrepository created. Alert routing is separated from metrics storage on day one — a decision that still shapes every alerting pipeline built on the stack. - 2015-01 — Public announcement via the SoundCloud engineering blog ("Prometheus: Monitoring at SoundCloud", Julius Volz & Björn Rabenstein). Boxever and Docker users adopt before the announcement even lands.
- 2016-05-09 — CNCF TOC votes unanimously to host Prometheus as the foundation's second project, after Kubernetes. The press release quotes TOC chair Alexis Richardson on the "new cloud native paradigm."
- 2016-07-18 — v1.0.0 published. The query language is PromQL, the exposition format is text, and the storage engine is already showing the strain that will force a rewrite.
- 2017 — Fabian Reinartz (fabxc) designs the new TSDB — time-sharded blocks, per-block indexes, delta-of-delta compression, a write-ahead log — that becomes v2.0's storage engine. The project has committed to throwing away its entire v1 storage layer rather than patching it.
- 2017-11-01 — The Thanos repository is created at Improbable by Fabian Reinartz and Bartłomiej Płotka; it is publicly open-sourced in June 2018. The first major attempt to bolt horizontal scale onto Prometheus via the v2 block format and object storage. More below.
- 2017-11-08 — v2.0.0 published: the TSDB rewrite ships. ~3x CPU reduction, ~2x disk-space reduction, ~100x I/O reduction (GrafanaCon slides, Tom Wilkie). The single most consequential engineering decision in the project's history — and the one that made the Thanos/Mimir ecosystems possible.
- 2017–2018 — OpenMetrics is spun out of Prometheus (repo created ) as an attempt to standardize the exposition format as a vendor-neutral IETF-track spec — and to separate "the Prometheus format" from "the Prometheus project."
- 2018-08-09 — CNCF announces Prometheus graduation at PromCon — the second project to graduate, after Kubernetes. Roughly 20 active maintainers, 1,000+ contributors, 13,000+ commits at the time.
- 2022-03 — Grafana Labs forks Cortex into Mimir, explicitly to "chip away at five years of accumulated technical debt" in the Cortex codebase — Grafana's own blog announcement uses those words. Cortex is effectively absorbed into the Grafana ecosystem, and Grafana Cloud is now the commercial center of gravity for the Prometheus ecosystem.
- 2022-11-08 — v2.40.0 ships experimental native histograms (
--enable-feature=native-histograms, #11447): sparse, high-resolution, no fixed buckets. The biggest data-model change since the TSDB rewrite. - 2024-07 — OpenMetrics archived by CNCF TOC action, merged back into Prometheus governance. The spec leaves home as an independent project and returns as a Prometheus working group — "OpenMetrics is dead, long live OpenMetrics (as Prometheus format)."
- 2024-11-14 — v3.0.0 published: first new major in seven years. New UI, UTF-8 metric names by default, Agent mode stable, auto-GOMEMLIMIT/GOMAXPROCS, OTLP receiver promoted, and a pile of removed deprecated flags whose migration guide is required reading.
- 2026-08-27 — PR #18606 ("model/textparse: implement OM2 scrape format", rbizos) merges. The parser that will become the OM2.0 scrape support in v3.15.0 lands in main.
- 2026-09-25 — v3.15.0 published — six days before this tale. OM2.0 scrape format implemented (#18606), Unix-domain-socket scraping merged (#12024, issue filed Feb 2023 by NitroCao — three and a half years to merge), zstd-compressed scrapes, XOR2 chunk encoding stabilized, runtime log-level reconfiguration, and Remote Write 2.0 era defaults.
Crisis Points & Architectural Pivots
1. The v2.0 storage rewrite (2017) — rewrite, don't patch
By 2016, Prometheus 1.x's storage was drowning under Kubernetes-grade label churn — "too many files, huge unsorted global index" is how the era's own talks describe it. The project's answer was not to patch but to replace: Fabian Reinartz's TSDB (time-sharded blocks, per-block indexes, Gorilla-style delta-of-delta compression, write-ahead log) shipped in v2.0.0 on with numbers that made the case for themselves — roughly 3x less CPU, 2x less disk, 100x less I/O per the GrafanaCon 2018 slides. LWN's coverage of the release called it "a hundred-fold I/O performance improvement." The v2 storage format is also what Thanos leverages — the Thanos project describes itself as leveraging "the Prometheus 2.0 storage format to cost-efficiently store historical metric data in any object storage." The rewrite didn't just fix Prometheus; it created the interface that the entire long-term-storage ecosystem would build on.
2. The great refusal: no clustering, ever
Prometheus has held the same line since its design doc era: one server, its own local disk, no built-in clustering or networked storage — HA by running two servers and deduplicating alerts in the Alertmanager, long-term storage via remote write to someone else's database. The project's own FAQ frames the alternative as "often more of a marketing claim than anything else" and says you can run a single instance "reliably with tens of millions of active series." This refusal is the single most consequential decision in the project's history — and the most productive. It is directly responsible for the entire long-term-storage ecosystem: Thanos (started 2017 at Improbable by Fabian Reinartz — the same engineer who wrote the TSDB — and Bartłomiej Płotka), Cortex (started 2016 by Julius Volz and Tom Wilkie), and Mimir (Grafana's 2022 Cortex fork, per the Mimir blog announcement). Three CNCF-track projects exist because Prometheus refused to become a database company.
3. The OpenMetrics round-trip (2017–2024)
In 2017 the exposition format was spun out as OpenMetrics — a CNCF sandbox→incubating project aiming for IETF standardization, created by Richard Hartmann and collaborators, and by 2022 accepted as a CNCF incubating project with an IETF Internet Draft underway. The goal was separation: "the Prometheus format" as a neutral standard independent of "the Prometheus project." Then reality arrived: two similar-but-divergent formats confused users, the IETF draft stalled, and the working group wind-down discussion ended with the CNCF TOC archiving the independent project and folding the spec back under Prometheus governance in July 2024. The repository moved to prometheus/OpenMetrics in October 2024, and a new OpenMetrics 2.0 working group was formed under Prometheus itself. The v3.15.0 release notes — six days old at time of writing — show where that round trip ends: "[FEATURE] scrape: Implement OM2.0 scrape format. #18606." The format that left home to become a standard has returned home as a feature flag.
4. The v3.0 deprecation purge (2024)
Seven years of accumulated feature flags had made Prometheus 2.x's CLI surface a museum. v3.0.0 (November 2024) removed --enable-feature=auto-gomemlimit, old-ui, expand-external-labels, the v1 Alertmanager API, and dozens more, while making UTF-8 metric names the default and shipping a brand-new UI. The migration guide grew so large it earned its own documentation page — one that, notably, 404s at the URL printed in the release notes themselves (the correct live path is under /docs/prometheus/latest/migration/, which we verified resolves). A major that is mostly subtraction is the most Prometheus move imaginable.
Community Engine & Corporate Influence: Who Actually Built It
The contributor table is a map of the cloud-native industry itself. Per the GitHub contributors API (October 2026):
CONTRIBUTIONS LOGIN WHO
1681 fabxc Fabian Reinartz — TSDB author; Thanos co-founder
1614 juliusv Julius Volz — SoundCloud; co-founder, long-time release manager
1069 beorn7 Björn Rabenstein — SoundCloud-era; PromQL and native histograms
994 bboreham Bryan Boreham
882 dependabot[bot] (automation)
633 roidelapluie Julien Pivotto — Prometheus team member
553 krajorama George Krajcsovits — Grafana
507 brian-brazil Brian Brazil — Robust Perception; PromQL spec, promtool
480 bwplotka Bartłomiej Płotka — Improbable→Grafana Labs; Thanos co-founder
347 codesome Ganesh Vernekar — Prometheus maintainer
324 matttproud Matt T. Proud — SoundCloud; author of the first commit
301 SuperQ Ben Kochie — Prometheus team
300 gouthamve Goutham Veeramachaneni
279 simonpasquier Simon Pasquier — Red Hat
Read the affiliations and the pattern is unmistakable: SoundCloud founded it, Weaveworks and Red Hat industrialized it, Grafana Labs commercialized the ecosystem around it, and one independent engineer (Fabian Reinartz) wrote the two highest-leverage pieces of code in the project's history — the TSDB, and then Thanos on top of it. No single company owns Prometheus; the graduation in 2018 moved stewardship to the CNCF, and that structure is why the project survived the death or absorption of so many early patrons — SoundCloud retrenched, Weaveworks shut down in early 2024, Improbable pivoted away from infrastructure. The foundation was not a formality; it was life insurance.
For platform teams, the commercial gravity today is easy to name. Grafana Labs operates the managed-Prometheus lineage — Mimir, forked from Cortex in 2022, and the Grafana Cloud metrics pipeline — and employs maintainers across the ecosystem; the CNCF project page lists the governance facts; and the project's own documentation remains the starting point for every self-hosted deployment.
Current Trajectory: The Verdict
Six days before this tale was written, Prometheus shipped v3.15.0. Its release notes read like a project settling its own history: OM2.0 scrape support (the format spec lives with the exposition format documentation) closes the OpenMetrics loop; Unix-domain-socket scraping closes a request filed three and a half years earlier; XOR2 chunk encoding stabilizes; zstd-compressed scrapes arrive behind a feature flag; runtime log-level reconfiguration lands. The cadence is the tell — v3.14.0 landed August 18, 2026, v3.15.0 on September 25. That is a healthy, boring, deliberate cadence for a fourteen-year-old codebase, and "boring" is the correct verdict for monitoring infrastructure.
The competitive pressure is real — OpenTelemetry's collector is the ingestion layer many teams now standardize on, and VictoriaMetrics built a business on drop-in remote-write compatibility (see our VictoriaMetrics vs. Prometheus review and the v3.15.0 release page for our current thinking) — but the fundamentals have not moved: pull-based scraping with explicit targets, PromQL, and single-node simplicity remain the default answer for a platform team's first serious metrics stack. The monoculture has held because the refusal at its center — no clustering, no storage product, just excellent single-node monitoring — turned every would-be competitor into a downstream component. Fourteen years in, Prometheus is not the future of observability. It is the foundation under it.