Prometheus 3.15: The Silent-Corruption Patch, a 3.5-Year Socket Ask, and the XOR2 One-Way Door
Sources
- Prometheus 3.15.0 release notes
- PR #19700 — Agent mode: reset WAL buffers after write errors
- PR #18091 — Scrape targets via Unix Domain Sockets
- Issue #12024 — Unix socket for metrics address (opened Feb 2023)
- PR #18606 — OM2 scrape format parser
- PR #19461 — XOR2 float chunk encoding stabilized
- PR #19511 — Reloadable runtime.log_level
- PR #18843 — auto-gomemlimit refresh interval
- PR #19502 — zstd-compressed scrape responses
- Prometheus configuration reference
- OpenMetrics 2.0 specification (experimental)
- OpenMetrics 2.0 migration guide for client libraries
- Thanos Sidecar documentation
- CRI-O PR #10253 — Unix-socket-only metrics endpoint
- Docker daemon Prometheus metrics docs
Prometheus 3.15.0 went out on September 25, 2026, and almost nobody in the CNCF ecosystem noticed the release that matters most to them. That is not marketing speak — it is the actual risk profile of this release. The changelog reads as a routine minor: a scrape-format addition, a long-awaited socket feature, a storage-encoding promotion. Buried in the bugfix list, though, sits PR #19700: a fix for a failure mode where a failed WAL write in Agent mode could return a dirty buffer to the pool, and a later append could concatenate a new record with the corpse of the failed one — producing a WAL record that passes its checksum and is structurally invalid. That is the quietest data-corruption class there is: no crash, no log line at write time, just a segment your future self cannot replay. This is the release analysis for people who run Prometheus in anger, not people who read feature lists.
The Bug Nobody Upgrades For: Agent-Mode WAL Corruption
Agent mode — the scrape-and-forward configuration where Prometheus keeps no local query path and streams samples via remote write — lives and dies by its write-ahead log. Every sample batch is appended to the WAL before it is shipped; on restart, the WAL replay is the difference between a blip and a data-loss incident. The bug fixed by PR #19700 (authored by @fancive, merged September 16, days before the release cut) is in the error path of that append, which is exactly where nobody looks until a disk fills up:
Normal append: log(rec) ─▶ WAL page write OK ─▶ buffer(len=N) ─▶ back to pool
Failed append: log(rec) ─▶ WAL page write FAILS
│
pre-3.15: └─▶ buffer returned to the shared pool
STILL HOLDING the partially-encoded
failed record
│
next append reuses that buffer:
new record bytes CONCATENATED with
the failed record's remnant
│
result: a WAL record that
CHECKSUMS but is structurally
invalid on replay
3.15 (#19700): all error paths in log() and logSeries() return buffers
to the pool with ZERO length — reuse starts cleanRead that again: the failed write already surfaces an error to the caller, so monitoring shows a rejected batch and operators move on. The corruption is deferred — it detonates at the next replay, when the WAL reader hits a record whose bytes are internally consistent enough to checksum but semantically garbage. The PR body is explicit that this was reproduced deterministically (go test ./tsdb/agent -run '^TestAppenderBufferResetAfterWALWriteError$' -count=50) and equally explicit about what it does not fix: the separate partial-page WAL corruption tracked in issue #13027, plus a distinct failed-appender reuse problem the PR deliberately leaves out of scope. If you run Agent mode, this fix alone is the reason to schedule the 3.15.0 rollout — do not wait for a CVE number that will never come, because this is not a security bug, it is a durability bug, and durability bugs never get a marketing push.
The companion fix, PR #19814 (merged by @roidelapluie one week before the release), tells the other half of the story: the Agent WAL reader now ignores unknown record types. That is a rollback-enabling change — if you upgrade to 3.15.0, run it, and need to fall back, a WAL containing record types your older binary does not recognize no longer poisons the replay. It is the kind of change you only appreciate during an incident at 3 a.m., which is precisely the audience this release serves.
Unix Domain Socket Scraping: Issue #12024, Finally Closed
The oldest debt paid in this release dates to February 27, 2023: issue #12024, the request to scrape metrics endpoints exposed over a Unix domain socket instead of TCP. Three and a half years of "me too" comments later, PR #18091 (authored by @IngmarStein, merged by @bwplotka on August 6) shipped it. The mechanics matter because they are not what you would guess: Prometheus does not build a unix:// URL. The target URL is constructed from __address__ and __scheme__ exactly as before — the socket path travels in a dedicated __unix_socket__ label (constant defined in scrape/scrape.go, line 2484) and is handed to the dialer through the request context. If __address__ is empty, the host falls back to localhost purely to keep the URL valid; the connection never touches TCP.
scrape_configs:
- job_name: envoy-admin
metrics_path: /stats/prometheus
static_configs:
- targets:
- envoy.internal # placeholder host; real connection dials the socket
labels:
__scheme__: http
__unix_socket__: /var/run/envoy/admin.sock
# relabel_configs can also set __unix_socket__ per target from any
# service-discovery label (__meta_*) — same rule as __address__.curl --unix-socket /var/run/envoy/admin.sock http://localhost/stats/prometheus | head
# works the same way over TLS if the exporter serves https on the socketThe honest adoption note: the feature is only as useful as the exporter ecosystem allows, and it is uneven. Envoy's admin interface is the canonical fit — its address accepts a pipe (Unix Domain Socket) instead of a socket address, it serves Prometheus output at /stats/prometheus, and until 3.15 scraping it required a TCP sidecar or curl --unix-socket in a textfile collector. CRI-O has an open PR to allow a socket-only metrics endpoint (metrics_socket, TCP listener disabled) — not merged at the time of writing, so CRI-O's metrics still require the TCP port today. Docker's daemon metrics endpoint is likewise still TCP-only (metrics-addr on loopback), with no Unix-socket option on its documented path. So the practical 3.15.0 win is targeted: Envoy-admin fleets and socket-first exporters where the operator's goal is to never bind a port at all. For everyone scraping over ordinary service discovery, this changes nothing — which is fine; the issue was never about you.
OM 2.0 Scraping: The Format That Ends the Great Metric-Name Renaming
The second headline feature is parser support for OpenMetrics 2.0 (PR #18606, authored by @rbizos, merged by @krajorama — the parser lives at model/textparse/openmetrics2parse.go in the release tree). OM 2.0 is the exposition format that stops pretending metric names must look like process_cpu_seconds_total. The migration guide frames it as the OTel-bridge peace treaty: OpenTelemetry-instrumented services emit dotted, UTF-8 names like process.cpu.seconds, and today every bridge mangles them into legacy-safe underscores. OM 2.0 lets the wire format carry the real name, quoted when needed.
Three more changes ride along with the name grammar, and for a scraping tier they are the ones that reduce cardinality and correctness drift: start timestamps move inline onto the sample line (st@) instead of requiring the separate synthetic _created samples; histograms collapse into a single CompositeValue block instead of the _bucket/_count/_sum triple; and exemplars switch to W3C trace-context keys (trace_id, span_id) instead of vendor-shaped key=value pairs. The negotiation is standard content negotiation with a twist worth knowing: OM 2.0 is UTF-8-native, so its Accept entry carries no escaping parameter at all — and it is not in the default protocol preference list. You opt in per scrape config:
scrape_configs:
- job_name: otel-services
scrape_protocols:
- OpenMetricsText2.0.0 # UTF-8 names, st@ timestamps, CompositeValue
- OpenMetricsText1.0.0 # fallback for endpoints not yet OM2-capable
- PrometheusText1.0.0
static_configs:
- targets: ["otel-gateway.internal:8888"]One documentation gap to be aware of, because it will bite someone doing config review: the generated configuration reference's comment listing supported scrape_protocols values still omits OpenMetricsText2.0.0, but the value is fully accepted — it is defined in config/config.go with its own application/openmetrics-text;version=2.0.0 media type and passes validation like any other protocol. The spec page itself is marked [EXPERIMENTAL], and the release notes' own guidance is to treat OM 2.0 adoption as an exporter-by-exporter rollout, not a global flip — exactly the caution you want from the project that also wrote the original OpenMetrics 1.0 spec. If your organization is heavy on OpenTelemetry collectors — and most platform teams now are — this is the first Prometheus release where the metric names in your dashboards can match the names in your application code.
XOR2 Becomes Stable — and Becomes a One-Way Door
The storage change is the one with career consequences. PR #19461 (authored and merged by @roidelapluie) stabilizes the XOR2 float chunk encoding and re-homes its configuration: the --enable-feature=xor2-encoding flag is deprecated and becomes a no-op in a future major version, replaced by a runtime-reloadable config field, storage.tsdb.chunk_encoding.floats: xor2. XOR2 gives better disk compression than XOR on typical workloads and can store start timestamps — which is why the st-storage feature flag now auto-enables XOR2 (and the ST-capable histogram encoding) rather than making you pass three flags in coordination (PR #19518).
Now the part the changelog buries in a WARNING comment: XOR2 chunks are a one-way door. Older Prometheus versions cannot read them. Downstream tools that read the TSDB directly cannot read them — the docs name the Thanos Sidecar's block uploads explicitly as the known breakage. And once XOR2 chunks exist on disk, the documented downgrade procedure is to delete the affected blocks from disk manually — there is no conversion path. Here is the decision layout:
enable storage.tsdb.chunk_encoding.floats: xor2
│
┌──────────────────────────────┼────────────────────────────┐
▼ ▼ ▼
Prometheus 3.15 Thanos / Cortex / LTS sidecars rollback to 3.14-:
queries: OK reading TSDB blocks directly: queries touching
(3.15 reads XOR "unknown encoding" errors until XOR2 blocks fail;
and XOR chunks alike) the downstream tool adds XOR2 documented fix is
support — sidecar block uploads deleting affected
are the documented breakage blocks from disk
BY HAND
│ │ │
└──── safe only when every reader of your blocks has ────────┘
committed to a XOR2-capable release, and you accept
that downgrade requires surgery, not a binary swapstorage:
tsdb:
chunk_encoding:
floats: xor2 # 'xor' (default) or 'xor2'; reloadable at runtime
# NOTE (from the docs): setting 'xor' alongside --enable-feature=
# st-storage is rejected at start/reload — XOR chunks cannot store
# start timestamps. The encoding only cuts over when the current
# chunk is next cut (size, time range, or sample count).Our operational guidance, formed by watching a decade of storage-format transitions: treat this as a fleet decision, not a config tweak. If you run plain Prometheus and read blocks only with your own Prometheus, XOR2 is a straightforward compression win and you should enable it on your next maintenance window. If anything else reads your TSDB path — a Thanos sidecar uploading to object storage, a backup tool, an LTS vendor — confirm XOR2 support on their current release before you flip it, because your rollback story is not "redeploy the old version", it is "identify and delete poisoned blocks while your retention window burns". That asymmetry — easy forward, surgical backward — is the single most important fact in this release.
The Quality-of-Life Pile, and One Deprecation to Schedule
Under the headlines, 3.15.0 carries a dense set of changes that platform teams will feel weekly. The ones worth your attention:
| Change | What actually happens | Why you care |
|---|---|---|
| Reloadable log level (PR #19511) | New runtime.log_level config field, reloadable live. The level only flips after every component reloader succeeds — a parse or apply failure keeps the previous level. --log.level is deprecated. | Debug-level logging on a production Prometheus without a restart — and without the foot-gun of a half-applied reload. |
--auto-gomemlimit.refresh-interval(PR #18843) | Periodically re-detects the container/system memory limit and updates GOMEMLIMIT at runtime. Default 0s keeps the old startup-only behavior. The help text warns: a downward limit change causes a temporary increase in GC activity. | Vertical Pod Autoscaler users: GOMEMLIMIT now tracks the pod's actual limit instead of the value at boot. This was the gap that made auto-GOMEMLIMIT and VPA quietly incompatible. |
| zstd scrape compression (PR #19502) | --enable-feature=zstd-scrape advertises zstd alongside gzip in Accept-Encoding. body_size_limit applies to the decompressed body. A target answering with zstd when it was not advertised fails the scrape. | Large exporters on slow links — think 100k-series kubelet endpoints over WAN edges. zstd's ratio/speed curve beats gzip for scrape payloads. |
| Head init busy-wait fix (PR #18001) | The TSDB head initialization busy-wait loop is gone; head-init CPU utilization drops. | Restart-heavy environments (bursty autoscaling, node recycling) stop paying a CPU tax on every WAL replay start. |
| Subquery sample-accounting fix (#18598) | Unaligned range-query ends no longer let subqueries evaluate past the parent's last step — which previously inflated peakSamples against query.max-samples and wasted I/O reading samples the result never used. | Fewer spurious "query processing would load too many samples" errors on otherwise fine queries — the class of false positive that erodes trust in the limits you set. |
On deprecations, put two dates in your calendar now: --log.level will go away in a future major, so move log-level management into the config file's runtime: block this quarter; and the xor2-encoding feature flag becomes a no-op in the next major, so anyone who enabled XOR2 via the flag should migrate to storage.tsdb.chunk_encoding.floats before then — same behavior, but the config field survives reloads and documents itself in your Prometheus config review instead of living in a Helm values appendix.
Hidden Costs and Breaking Changes
- The Agent-mode fixes are the quiet headline. If you run Prometheus in agent mode, the WAL buffer-reset fix (#19700) and the unknown-record-type tolerance (#19814) are durability and rollback guarantees, respectively. Plan the upgrade around them, not around the feature list.
- XOR2 is a coordination event, not a flag flip. Every direct reader of your TSDB blocks must be XOR2-capable before you enable it, and your rollback plan must name the block-deletion procedure. If you cannot write that plan in one paragraph, you are not ready to enable it.
- OM 2.0 is opt-in per scrape job and experimental. The spec page carries the [EXPERIMENTAL] label and the default protocol preference list does not include it — adopt it exporter-by-exporter and keep an OM 1.0 fallback in the negotiation list, as the release notes themselves recommend.
- The configuration reference comment lags the code on
scrape_protocolsaccepted values (OpenMetricsText2.0.0is valid but unlisted in the comment). Do not let a config-review checklist reject valid OM 2.0 configs. - UDS scraping changes no defaults. Nothing moves until you set
__unix_socket__on a target; existing TCP scrape paths are untouched, including theunix:address form discussed in the original PR thread — the shipped mechanism is the__unix_socket__label.
Verdict
Skip this release only if you run no Prometheus at all. Everyone else gets something real: agent-mode operators get a silent-corruption fix and a rollback guarantee they should not wait on; VPA users get GOMEMLIMIT that tracks reality; teams fighting query.max-samples false positives get honest sample accounting; and anyone with socket-only exporters (CRI-O's metrics_socket pattern) can finally drop the localhost-port workaround. The one decision in this release that deserves a design doc rather than a config commit is XOR2 — it is a storage-format commitment whose rollback is manual block deletion, coordinated with every downstream block reader you run. Enable it deliberately, with your Thanos sidecar and LTS tooling versions written down, or do not enable it yet. And if you want the operational posture this release quietly assumes — someone watching WAL health, negotiating scrape protocols deliberately, treating storage encodings as contracts — that is a platform-engineering function, not a side task. Managed options exist, from Grafana Cloud's hosted Prometheus to Amazon Managed Service for Prometheus; the durability bugs are theirs to catch, which is worth exactly what it costs you.
Credit Where Due
The release was assembled by github-actions but written by people: @fancive (zstd scrape compression and the WAL buffer-reset fix — two of the most operationally important changes in the release), @IngmarStein for finally landing the 3.5-year-old Unix-socket ask (filed by @NitroCao in February 2023, kept alive by commenters @DaAwesomeP, @h7x4, and @maxant through three summers), @rbizos for the OM 2.0 parser, @roidelapluie for the XOR2 stabilization and the agent WAL tolerance, @loglapa for the reloadable log level (closing issue #10352), and @sanidhyasin for the GOMEMLIMIT refresh interval. Mergers @bwplotka and @krajorama carried the review load. That is a healthy distribution: no single hero, several first-time-feeling contributions, and one very old debt paid in full.