Prometheus: The Monitoring Monoculture That Refused to Cluster

Ex-Googlers at SoundCloud, missing Borgmon, built the pull-based monitoring monoculture that graduated second from the CNCF — and stayed single-node on purpose.

Sources

Every ecosystem has one tool whose name became the category. For cloud-native metrics, that tool is Prometheus — and its story is the story of an entire industry learning to monitor what containers broke. It starts with two SoundCloud engineers missing an internal Google system they could no longer use, and ends — for now — with a fourteen-year-old codebase shipping its own wire format into the spec that was once spun out to escape it.

The Origin: Borgmon Withdrawal at SoundCloud

Ex-Googlers, StatsD pain, and a 2012 pet project

The design DNA of Prometheus is not original — it is Borgmon, Google's internal time-series monitoring system, described publicly in the SRE Book's chapter on practical alerting. Matt T. Proud and Julius Volz had both worked at Google and left. At SoundCloud, the pair hit the wall every fast-growing microservice shop hit in 2012: hundreds of services, thousands of instances churning constantly, and a StatsD+Graphite stack that, in their own words, ran into "a number of serious limitations." The founding SoundCloud blog post — written by Julius Volz and Björn Rabenstein in January 2015 — describes the wish list exactly: multi-dimensional data model, operational simplicity, decentralized collection, and a powerful query language, all in one system. Their conclusion: "we could not identify a system that combined them all until a colleague started an ambitious pet project in 2012 that aimed to do so."

The git archaeology confirms the story. The repository was created on , and the deepest commit page reads:

# Deepest page of prometheus/prometheus commit history:
$ gh api "repos/prometheus/prometheus/commits?until=2012-12-31T23:59:59Z&per_page=100"

DATE       SHA         AUTHOR          MESSAGE
2012-11-24 734d28b515 Matt T. Proud   Initial commit
2012-12-09 577acf4fe7 Matt T. Proud   Exploding the storage infrastructure by contexts
2012-12-11 6589fc92f8 Matt T. Proud   Strip web services, which weren't adding value
2012-12-12 59a708f25a Matt T. Proud   Provide prototype of storage layer interfaces
2012-12-12 0886592ebc Matt T. Proud   New interface definition after discussion
2012-12-19 a14dbd5bd0 Matt T. Proud   Interim commit for Julius

# Every commit on that first page: Matt T. Proud, prototyping interfaces,
# and handing interim state to Julius Volz. A pet project, in public.

SoundCloud put it into production by 2013 — before any announcement — and the alertmanager repository followed on , split out early because alerting is a different lifecycle problem from collection. The project went public in January 2015 with the SoundCloud blog post above, the first PromCon followed, and the rest is landscape history: in May 2016 the CNCF's TOC voted unanimously to accept Prometheus as the foundation's second hosted project after Kubernetes — and in August 2018, at PromCon, it became the second project to graduate, after Kubernetes. Not bad for a monitoring tool. Not bad at all.

The Timeline: From Pet Project to Graduated Standard

Crisis Points & Architectural Pivots

1. The v2.0 storage rewrite (2017) — rewrite, don't patch

By 2016, Prometheus 1.x's storage was drowning under Kubernetes-grade label churn — "too many files, huge unsorted global index" is how the era's own talks describe it. The project's answer was not to patch but to replace: Fabian Reinartz's TSDB (time-sharded blocks, per-block indexes, Gorilla-style delta-of-delta compression, write-ahead log) shipped in v2.0.0 on with numbers that made the case for themselves — roughly 3x less CPU, 2x less disk, 100x less I/O per the GrafanaCon 2018 slides. LWN's coverage of the release called it "a hundred-fold I/O performance improvement." The v2 storage format is also what Thanos leverages — the Thanos project describes itself as leveraging "the Prometheus 2.0 storage format to cost-efficiently store historical metric data in any object storage." The rewrite didn't just fix Prometheus; it created the interface that the entire long-term-storage ecosystem would build on.

2. The great refusal: no clustering, ever

Prometheus has held the same line since its design doc era: one server, its own local disk, no built-in clustering or networked storage — HA by running two servers and deduplicating alerts in the Alertmanager, long-term storage via remote write to someone else's database. The project's own FAQ frames the alternative as "often more of a marketing claim than anything else" and says you can run a single instance "reliably with tens of millions of active series." This refusal is the single most consequential decision in the project's history — and the most productive. It is directly responsible for the entire long-term-storage ecosystem: Thanos (started 2017 at Improbable by Fabian Reinartz — the same engineer who wrote the TSDB — and Bartłomiej Płotka), Cortex (started 2016 by Julius Volz and Tom Wilkie), and Mimir (Grafana's 2022 Cortex fork, per the Mimir blog announcement). Three CNCF-track projects exist because Prometheus refused to become a database company.

3. The OpenMetrics round-trip (2017–2024)

In 2017 the exposition format was spun out as OpenMetrics — a CNCF sandbox→incubating project aiming for IETF standardization, created by Richard Hartmann and collaborators, and by 2022 accepted as a CNCF incubating project with an IETF Internet Draft underway. The goal was separation: "the Prometheus format" as a neutral standard independent of "the Prometheus project." Then reality arrived: two similar-but-divergent formats confused users, the IETF draft stalled, and the working group wind-down discussion ended with the CNCF TOC archiving the independent project and folding the spec back under Prometheus governance in July 2024. The repository moved to prometheus/OpenMetrics in October 2024, and a new OpenMetrics 2.0 working group was formed under Prometheus itself. The v3.15.0 release notes — six days old at time of writing — show where that round trip ends: "[FEATURE] scrape: Implement OM2.0 scrape format. #18606." The format that left home to become a standard has returned home as a feature flag.

4. The v3.0 deprecation purge (2024)

Seven years of accumulated feature flags had made Prometheus 2.x's CLI surface a museum. v3.0.0 (November 2024) removed --enable-feature=auto-gomemlimit, old-ui, expand-external-labels, the v1 Alertmanager API, and dozens more, while making UTF-8 metric names the default and shipping a brand-new UI. The migration guide grew so large it earned its own documentation page — one that, notably, 404s at the URL printed in the release notes themselves (the correct live path is under /docs/prometheus/latest/migration/, which we verified resolves). A major that is mostly subtraction is the most Prometheus move imaginable.

Community Engine & Corporate Influence: Who Actually Built It

The contributor table is a map of the cloud-native industry itself. Per the GitHub contributors API (October 2026):

CONTRIBUTIONS  LOGIN            WHO
1681          fabxc            Fabian Reinartz — TSDB author; Thanos co-founder
1614          juliusv          Julius Volz — SoundCloud; co-founder, long-time release manager
1069          beorn7           Björn Rabenstein — SoundCloud-era; PromQL and native histograms
 994          bboreham         Bryan Boreham
 882          dependabot[bot]  (automation)
 633          roidelapluie     Julien Pivotto — Prometheus team member
 553          krajorama        George Krajcsovits — Grafana
 507          brian-brazil     Brian Brazil — Robust Perception; PromQL spec, promtool
 480          bwplotka         Bartłomiej Płotka — Improbable→Grafana Labs; Thanos co-founder
 347          codesome         Ganesh Vernekar — Prometheus maintainer
 324          matttproud       Matt T. Proud — SoundCloud; author of the first commit
 301          SuperQ           Ben Kochie — Prometheus team
 300          gouthamve        Goutham Veeramachaneni
 279          simonpasquier    Simon Pasquier — Red Hat

Read the affiliations and the pattern is unmistakable: SoundCloud founded it, Weaveworks and Red Hat industrialized it, Grafana Labs commercialized the ecosystem around it, and one independent engineer (Fabian Reinartz) wrote the two highest-leverage pieces of code in the project's history — the TSDB, and then Thanos on top of it. No single company owns Prometheus; the graduation in 2018 moved stewardship to the CNCF, and that structure is why the project survived the death or absorption of so many early patrons — SoundCloud retrenched, Weaveworks shut down in early 2024, Improbable pivoted away from infrastructure. The foundation was not a formality; it was life insurance.

For platform teams, the commercial gravity today is easy to name. Grafana Labs operates the managed-Prometheus lineage — Mimir, forked from Cortex in 2022, and the Grafana Cloud metrics pipeline — and employs maintainers across the ecosystem; the CNCF project page lists the governance facts; and the project's own documentation remains the starting point for every self-hosted deployment.

Current Trajectory: The Verdict

Six days before this tale was written, Prometheus shipped v3.15.0. Its release notes read like a project settling its own history: OM2.0 scrape support (the format spec lives with the exposition format documentation) closes the OpenMetrics loop; Unix-domain-socket scraping closes a request filed three and a half years earlier; XOR2 chunk encoding stabilizes; zstd-compressed scrapes arrive behind a feature flag; runtime log-level reconfiguration lands. The cadence is the tell — v3.14.0 landed August 18, 2026, v3.15.0 on September 25. That is a healthy, boring, deliberate cadence for a fourteen-year-old codebase, and "boring" is the correct verdict for monitoring infrastructure.

The competitive pressure is real — OpenTelemetry's collector is the ingestion layer many teams now standardize on, and VictoriaMetrics built a business on drop-in remote-write compatibility (see our VictoriaMetrics vs. Prometheus review and the v3.15.0 release page for our current thinking) — but the fundamentals have not moved: pull-based scraping with explicit targets, PromQL, and single-node simplicity remain the default answer for a platform team's first serious metrics stack. The monoculture has held because the refusal at its center — no clustering, no storage product, just excellent single-node monitoring — turned every would-be competitor into a downstream component. Fourteen years in, Prometheus is not the future of observability. It is the foundation under it.