VictoriaMetrics Review: Blazing Ingest, a 30-Second Blind Spot, and a Delete Button Without a Password
Sources
- VictoriaMetrics v1.153.0 release notes
- VictoriaMetrics single-node documentation
- VictoriaMetrics cluster documentation
- vmagent documentation (persistent queues, remote write)
- vmauth documentation (OIDC SSO, new in v1.153.0)
- MetricsQL documentation
- VictoriaMetrics Enterprise features
- Prometheus storage documentation (local TSDB limits)
- vmctl migration tool
- VictoriaMetrics GitHub repository (Apache-2.0)
Every platform team running vanilla Prometheus eventually hits the same wall: the local TSDB is a single-node database, retention fights with disk space, and the "just add Thanos" answer adds four more services than anyone wanted. VictoriaMetrics — the Apache-2.0 TSDB from the Latvia-based company of the same name — is the most credible "just replace the storage" answer in the CNCF-adjacent market, and this week it shipped v1.153.0 with two security advisories, OIDC single sign-on, and a batch of operational fixes. The repository turned eight years old the same week — created 2018-09-30, now at 17,788 stars and a release cadence that lands a minor every two weeks.
We didn't read about this one: we downloaded the actual v1.153.0 linux-amd64 binary, verified its SHA-256 checksum against the published checksums file, and ran it through a 300,000-sample high-churn ingest, restart cycles, Unix-socket permission tests, and a deliberate probe of every admin endpoint. The result is a two-sided verdict, like most honest ones: the storage engine is genuinely excellent, the release hygiene is better than most vendors', and the operational defaults contain at least three traps that will bite anyone who deploys it the way the quickstart suggests.
Executive Scorecard
| Dimension | Score | Why |
|---|---|---|
| Reliability | 8/10 | 0.2s clean shutdown, 0.2s storage reopen with all 503,613 rows intact, per-URL persistent disk queues in vmagent; monthly partitions make retention pruning a non-event. Loses points for a 5-second unflushed-data window and single-node being a single point of failure by design |
| DX | 7/10 | Drop-in Prometheus remote-write target, MetricsQL is a strict superset, vmui is genuinely useful; but ~1,000 flags, a two-stage data-visibility model (5s flush + 30s latency offset) that looks like data loss until you read the docs, and dashboards that silently show nothing for fresh series |
| Cost | 9/10 community / 4/10 enterprise | Community edition is free and includes the full storage engine; downsampling, multiple retentions, and automatic vmstorage discovery are Enterprise-gated behind a quote-only plans page, or a managed VictoriaMetrics Cloud |
| Security | 5/10 | v1.153.0 fixes a memory-exhaustion advisory in the native import path and an OIDC redirect flaw, and adds OIDC SSO — but /api/v1/admin/tsdb/delete_series accepts unauthenticated deletes out of the box, and the docs' own security section assumes you'll front everything with vmauth yourself |
Scope note, the honest kind: all numbers below were measured on this review's host — 4 CPU cores, 30 GB RAM, local SSD, single-node binary, default flags unless stated. They are single-run wall-clock numbers from a working benchmark, not a normalized multi-host benchmark suite; treat them as directional, and re-run on your own hardware before making a procurement decision. What they reliably demonstrate is the shape of the system: where it's fast, where it's clever, and where it will surprise your on-call engineers.
Architecture Mechanics: What Actually Runs
The single-node binary is one process with one job: absorb samples, index them, keep queries off the disk where possible. The design choices that matter operationally are all visible in the startup log:
remote writers (Prometheus, vmagent, OTel, Influx, Datadog agent)
|
| /api/v1/write (snappy+protobuf remote write)
| /api/v1/import/prometheus (exposition text, fastest bulk path)
| /api/v1/import (JSON line)
| /api/v1/import/csv, /influx/write, /opentelemetry/v1/metrics
v
+-----------------------------------------------------+
| victoria-metrics-prod (single process) |
| |
| in-memory rows --flush every 5s--> data parts |
| (-inmemoryDataFlushInterval, min 1s) |
| |
| storage layout (per -storageDataPath): |
| data/data/small/YYYY_MM small parts |
| data/data/big/YYYY_MM big parts (compacted) |
| data/indexdb/YYYY_MM inverted index |
| cache/rollupResult query result cache |
| snapshots/, metadata/, tmp/ |
| |
| monthly partitions: retention prune = drop month |
| adaptive limits at boot: |
| -search.maxUniqueTimeseries derived from |
| remaining RAM (925,535 on this 30GB host) |
+-----------------------------------------------------+
|
| PromQL/MetricsQL queries, -search.latencyOffset=30s
| (samples < 30s old excluded by default)
v
/api/v1/query, /api/v1/query_range, /api/v1/series,
/api/v1/export, Graphite API, vmui (web UI)
Three mechanics from that diagram drive everything else in this review. First, ingestion is decoupled from persistence: samples sit in memory for up to 5 seconds (-inmemoryDataFlushInterval, minimum 1s) before a guaranteed disk flush — the docs explicitly note the flush is what survives unclean shutdowns like OOM kills and power loss. Second, query results exclude the last 30 seconds of samples by default (-search.latencyOffset, overridable per query via a latency_offset argument, floor 1ms) — we verify both windows live below, because their combination is the number-one "VictoriaMetrics ate my data" false alarm. Third, data is organized into monthly partitions: on our test host the boot log showed partition "2026_09" has been created and creating a partition "2026_10" — which means retention is enforced by dropping whole months, the same trick ClickHouse users will recognize, and it makes -retentionPeriod=100y-style configs operationally boring (our test ran exactly that).
The cluster version fans the same storage engine out into three services — vmstorage (state), vminsert (write router), vmselect (query router) — with multitenancy as a first-class URL path (/insert/), replication via -replicationFactor at vminsert, and per-tenant limits in the Enterprise edition. The single-node binary does not support multitenancy at all: the docs are explicit that when single-node is read through -vmselectAddr, everything belongs to the 0:0 tenant. Teams that need hard tenant isolation on the write path are in cluster territory, full stop.
Where does Prometheus fit? VictoriaMetrics' own guidance — and our testing backs the shape of it — is that the single-node binary comfortably handles ingestion below a million samples per second, while Prometheus' local TSDB remains the right scraper for small estates: it is a single-node database in its own words ("not clustered or replicated... should be managed like any other single node database") with two-hour blocks and an explicit steer toward remote storage for durability. The realistic architectures are therefore: Prometheus scraping + VictoriaMetrics storing (remote write, the common case), vmagent scraping + VictoriaMetrics storing (replaces the Prometheus server entirely, adds per-URL persistent disk queues and relabeling), or VictoriaMetrics cluster (horizontal storage, multitenancy, and the Enterprise features). What does not work — and we've watched teams try — is treating single-node VictoriaMetrics as both scraper replacement and long-term system of record on one box, with no replication and no disk redundancy.
Hands-On: Ingest, Query, and the Two Windows That Look Like Data Loss
We ran the real binary end to end. Setup is exactly as advertised — untar, run with a data path and retention, health endpoint responds immediately:
$ sha256sum victoria-metrics-linux-amd64-v1.153.0.tar.gz
1b495bde563825cf83dc7c0425a9d8fa03e7214858e0949f9177efc6bb1f8bfc
$ ./victoria-metrics-prod -storageDataPath=/tmp/vmtest/data \
-retentionPeriod=100y -httpListenAddr=:8428
# startup log (verbatim, 4-core / 30GB host):
# partition "2026_09" has been created
# creating a partition "2026_10" with smallPartsPath=...
# limiting -search.maxUniqueTimeseries to 925535 according to
# -search.maxConcurrentRequests=8 and remaining memory=1480856372 bytes
# successfully opened storage in 0.107 seconds
# pprof handlers are exposed at http://0.0.0.0:8428/debug/pprof/
$ curl -s http://localhost:8428/health
OKThat pprof handlers are exposed line at startup is your first security clue: this is a component that assumes it lives on a trusted network segment unless you lock it down. More on that in the failure-modes section.
Ingest throughput: the headline claim holds
We generated a hostile payload — 5,000 series with pod labels that churn six times per series (30,000 distinct label sets), 300,000 total samples — and pushed it through the two main ingest paths. On this modest 4-core host:
| Ingest path | Wall time | Effective rate | Notes |
|---|---|---|---|
/api/v1/import/prometheus (exposition text, 32.7MB body) | 1.08s | ~277,000 samples/s | HTTP 204, zero errors. Docs correctly warn this path is "optimized for performance" and not for lossless streaming ingestion |
/api/v1/import (JSON line, same 300k samples) | 4.33s | ~69,000 samples/s | Same data, 4x slower — format choice is a real decision at scale |
~277k samples/s on four cores with default flags is consistent with the project's long-standing reputation here, and it's the single strongest argument for the migration: that is comfortably above what a mid-size fleet scrapes, on one process, with headroom to spare. After the load test the index reported totalSeries: 259,317 and totalLabelValuePairs: 1,555,879 — our label churn created a quarter-million live series, and neither ingest nor the subsequent full-index aggregation query (sum(rate(http_requests_total[5m])) by (dc) fetching 198,211 series in 629ms) degraded. A 5000-series rate() over the raw churn set took 2.7s — heavier, because per-series rollups over hostile label sets are index-bound, and that is where your capacity planning lives.
The two-window gotcha: 5 seconds of flush, 30 seconds of latency offset
Here is the trap we hit on day one, and the one that will generate most of your false "data is missing" incidents. We imported a sample and queried it back one second later: HTTP 204 on import, empty result on query. No error, no log line, no hint. The sample was in the process — /api/v1/export couldn't see it either. The mechanics, verified live:
$ # import one sample with an explicit timestamp, then poll:
$ curl -d 'lo_probe2{age="fresh"} 99' -X POST \
http://localhost:8428/api/v1/import/prometheus # HTTP 204
t+2s /api/v1/export -> empty <- flush window
t+2s /api/v1/query -> empty (latency_offset=1ms)
t+5s /api/v1/export -> shows sample <- 5s flush landed
t+5s /api/v1/query
?latency_offset=1ms -> shows sample <- offset override works
t+5s /api/v1/query -> empty <- default 30s offset
t+10s /api/v1/query -> empty <- still hidden
t+32s /api/v1/query -> shows sample <- default window expiresTwo independent mechanisms, both defaults: samples are not guaranteed on disk (or visible to export) until the 5-second -inmemoryDataFlushInterval fires — that flush is what survives an OOM kill, per the docs — and even after flush, instant queries exclude anything newer than -search.latencyOffset (30s) so that late, out-of-order samples from slow scrapers don't tear holes in the last data point. The per-query escape hatch is real but has a floor: latency_offset=0s returns 400: latency_offset=0ms is out of allowed range [1ms ... 3153600000000ms], while latency_offset=1ms works. Prometheus has no equivalent visibility gap for freshly scraped local data — a scraper reads its own head immediately — so teams migrating dashboards will see "missing" panels for the first 30 seconds of every series' life. Operational fix: either tune -search.latencyOffset down fleet-wide (at the documented cost of occasionally incomplete last points for slow/lagging writers), or bake latency_offset=1ms into internal "did the data land" probes — never into customer-facing alerting, where the guard exists for good reason.
Restart, durability, and the umask fix
Graceful shutdown and reopen is where the storage engine shows its quality: SIGTERM shut the process down in 0.203s with the rollup cache saved; restart reopened the same data directory in 0.215s with all partsCount: 6; blocksCount: 498650; rowsCount: 503613 intact and queryable. Data does not go through a risky block-compaction gauntlet on restart the way large Prometheus TSDBs can after an unclean kill.
One v1.153.0 bugfix earned specific verification because it's the kind of thing that breaks real deployments: Unix-domain-socket listeners previously hardcoded 0600 permissions, which meant a reverse proxy running as another user simply could not connect. The fix (#11615) makes socket permissions follow process umask. Confirmed live on v1.153.0:
$ umask 077 && ./victoria-metrics-prod -httpListenAddr=unix:/tmp/vm.sock ...
$ stat -c '%a %n' /tmp/vm.sock
700 /tmp/vm.sock # <- follows umask now
$ umask 022 && ./victoria-metrics-prod -httpListenAddr=unix:/tmp/vm2.sock ...
$ stat -c '%a %n' /tmp/vm2.sock
755 /tmp/vm2.sock # <- also follows umask (mind your umask!)
$ curl --unix-socket /tmp/vm.sock http://localhost/health
OKThe flip side: a careless umask now produces a world-connectable metrics socket. The footgun changed shape; it didn't disappear.
What v1.153.0 Actually Changed (and What It Signals)
Reading the release notes as an operator rather than a changelog consumer, three items stand out:
- Two security fixes, both worth understanding before you upgrade-or-ignore.
GHSA-8g4f-32hw-vqf8:vminsert,vmagent, andvmsinglenow correctly apply memory limits to zstd-encoded blocks arriving via the native import path — previously a maliciously crafted request could balloon memory allocation during ingestion. If your import endpoints are network-reachable (see the delete-series finding below), this is your patch-or-else item.GHSA-xxqh-2hcc-9fp6:vmauth's OIDC Discovery HTTP client no longer follows redirects outside the original request host. Both were also backported to v1.148.5 and v1.136.19 — the project maintains LTS branches, which is more release hygiene than most OSS vendors manage. (The advisory detail pages were not yet published at review time; the descriptions here come from the release notes and fix PR #11616.) - vmauth learned OIDC single sign-on (docs, #10278): unauthenticated browser requests redirect to your IdP, vmauth verifies the response and sets a session cookie. This is a real feature for the "self-hosted Grafana behind SSO" pattern — it removes the last common excuse for putting dashboards behind basic auth — but note it landed in vmauth, the router, not in the storage components themselves.
- Direct-to-vmstorage remote write (#11252, opt-in via
-enableIngestionAPI): a documented escape hatch that skips vminsert in the cluster. Read it as an acknowledgement that at extreme ingest rates the router can be the bottleneck — and as a trap, because bypassing vminsert also bypasses its replication and retry machinery.
Smaller but operationally pleasant: vmalert can now restore an alert directly to firing state after restart instead of sitting through a redundant pending period (#11401) — closing a genuine "we restarted the alerter and swallowed the incident" gap; vmalert actually reloads alert_relabel_configs on config reload instead of only at startup (#11635); empty extra_label/extra_filters[] query args are ignored instead of erroring (#11618) — which matters because it makes vmauth-based query-arg scrubbing work as documented; and stream aggregation reverted a parallelization that cost CPU for negligible latency gain (#9878). The pattern across the list is a project maturing: fewer new frontiers, more "we found the thing that was quietly wrong."
The Failure Modes You Will Actually Hit
1. The unauthenticated delete API (the finding that should change your deployment). With default flags, POST /api/v1/admin/tsdb/delete_series?match[]=... just works. We tested it: no auth header, no auth key, HTTP 204, series gone — then confirmed via export that the data was really deleted, and found the operation in the logs (Deleted 1 series). The same applies to /tags/delSeries. The fix is one flag — -deleteAuthKey=... forces an auth key for exactly these endpoints — plus, per the docs' own security section, fronting the whole HTTP surface with vmauth or an equivalent gateway, because admin endpoints, /debug/pprof, and /metrics are all exposed on the same listener otherwise. If VictoriaMetrics is reachable by anything other than a trusted service network, this is a data-destruction primitive sitting on port 8428. In 2026, "the network is the perimeter" is not a design, it's a regression.
2. High-churn label sets are the real capacity ceiling. Our synthetic pod-churn payload produced 259,317 live series from a conceptual 5,000 — because every distinct label set is a new series, and the index bills you for all of them (totalLabelValuePairs: 1,555,879 from one 300k-sample burst). VictoriaMetrics handles churn better than most (that is the core compression story), but the failure mode is silent: ingest stays fast, queries stay correct, and RAM use creeps with index size. The boot-time adaptive limit (-search.maxUniqueTimeseries was auto-derived to 925,535 on our 30GB host, explicitly tied to -search.maxConcurrentRequests=8 and remaining memory) will eventually refuse queries rather than OOM — which is the right failure mode, but you want to catch the trend on /api/v1/status/tsdb (excellent cardinality breakdown by metric name, label name, and label pair — a Grafana-billable feature shipped free) before the limit starts refusing queries during an incident. The Enterprise "per-tenant resource limits" page exists precisely because shared multitenant clusters without caps are where VictoriaMetrics deployments go to die.
3. Enterprise gating on exactly the retention-economics features. Here is the commercial sharp edge: community edition gives you one global -retentionPeriod. Downsampling (-downsampling.period=30d:5m,180d:1h) and multiple retentions per dataset are Enterprise features. The docs do offer the honest community workaround — run N instances with different -retentionPeriod + -storageDataPath behind vmauth routing — but that is now N databases to back up and monitor. The economics argument that sold VictoriaMetrics ("cheaper than Thanos at 13x compression") quietly assumes you don't need tiered retention, and most teams that keep metrics a year discover they do. Factor the Enterprise license or the managed VictoriaMetrics Cloud into the TCO from day one, because that is where the retention features live.
4. The single-node ceiling is real and arrives without warning. The docs' own line: single-node is recommended below a million samples per second and "can be set up in High Availability mode" — but HA here means two independent nodes and deduplicated remote write to both (via -dedup.minScrapeInterval), which is a read story, not a storage-consistency story. There is no replication, no shared-nothing cluster failover, and a dead node means its unflushed 5-second window (plus whatever the last vmagent queue hasn't replayed) is gone until you restore from vmbackup. The correct sizing question for single-node is not "how fast can it ingest" but "what is my RPO if this box dies" — and with default flags the honest answer is five seconds of writes plus your backup interval.
Migration Path, Honestly Assessed
The drop-in claim is mostly true: point Prometheus' remote_write at <vm>:8428/api/v1/write and Grafana stops being able to tell the difference, because the query API is Prometheus-compatible with MetricsQL as a strict superset. Two differences will surface in your dashboards, both verified from the docs and our run: increase()/rate() don't extrapolate (slow integer counters return integers, not Prometheus' fractional estimates — your SLO charts will move by a few percent and that is VictoriaMetrics being more correct, but a dashboard-diff will flag it); and metric names are dropped after functions/binary operators unless you opt in with keep_metric_names. For the data itself, vmctl block-migrates Prometheus TSDB snapshots into VictoriaMetrics natively. The vmagent swap is where teams save the most operational weight: it speaks Prometheus scrape_configs natively (-promscrape.config=/etc/prometheus/prometheus.yml, with -promscrape.config.strictParse=false to tolerate unsupported sections), and every -remoteWrite.url gets its own persistent on-disk queue that survives remote-storage outages and replays in FIFO order — with -remoteWrite.streamParse for the documented OOM protection on very high-cardinality scrapes. That is a genuinely better collector than raw Prometheus at fleet scale, and it's the piece of this stack we'd adopt first.
Verdict
VictoriaMetrics is the best default answer in open-source metrics storage, and v1.153.0 is a good version of it. The storage engine earns the reputation: quarter-million-series ingest at 277k samples/s on four cores, 629ms full-index aggregations, 0.2s clean shutdown and reopen with every row intact, monthly partitions that make retention boring, and a security posture that is improving in the right direction (memory-limit advisory fixed and LTS-backported, OIDC redirect contained, vmauth SSO added, umask sockets fixed). The Prometheus-compatibility surface is faithful enough that migration is an ops project, not a re-platform, and vmagent's persistent queues are strictly better than what raw Prometheus gives you at fleet scale.
But deploy it like the quickstart shows and you will get exactly three nasty surprises: panels that look empty for the first 30 seconds of every series' life (flush window + -search.latencyOffset), a metrics port where anyone on the network can permanently delete data (-deleteAuthKey unset), and the late discovery that the retention-economics features that justified the whole migration — downsampling, tiered retention — are behind the Enterprise license. None of these are hidden; all of them are documented. The gap between "documented" and "configured correctly in your environment" is where this review spent most of its time, and it is the actual adoption cost.
Who should skip VictoriaMetrics entirely
- Teams that need tiered retention or downsampling and won't buy Enterprise — the community workaround (N single-retention instances behind vmauth) multiplies your backup and monitoring surface. If 13x compression without downsampling already fits your budget, fine; if you're counting on downsampling for the long-term math, price the license first.
- Anyone unwilling to run a fronting gateway — until you put vmauth (or your own ingress) in front with auth and route filtering, you are one unauthenticated
POSTaway from a very bad day. This is a hard prerequisite, not an option. - Hard-multitenant SaaS on the write path — single-node has no multitenancy at all (everything is tenant
0:0); multitenancy means the cluster version, and per-tenant resource caps (the thing that keeps one noisy tenant from degrading everyone) are Enterprise. - Small estates already happy with vanilla Prometheus — if a single Prometheus with 15–30 days of local retention covers you and you have no long-term-storage requirement, you'd be adopting a second database to solve a problem you don't have. Revisit when retention pressure or federation actually arrives.
- Teams that need traces and logs in the same query plane — VictoriaLogs exists as a separate product with its own engine; the "one observability backend" story is looser than what Grafana Mimir-plus-Loki-plus-Tempo markets as an integrated stack. If vendor-stack cohesion matters more than per-engine excellence, weigh accordingly.
For everyone else — Prometheus estates drowning in disk, platform teams that want long retention without the Thanos/Cortex/Mimir sidecar sprawl, and anyone who has already accepted "metrics storage is a database and I will operate it like one" — VictoriaMetrics is the correct default, and the migration is cheaper than the quarter you'll spend avoiding it. Just front it with vmauth, set -deleteAuthKey before the first day, and decide your retention-tiering answer before the disk fills, not after.
References & Further Reading
- VictoriaMetrics v1.153.0 release notes — the source for every v1.153.0 change cited here (GHSA-8g4f-32hw-vqf8, GHSA-xxqh-2hcc-9fp6, OIDC SSO, #11252, #11401, #11615, #11618, #11635)
- VictoriaMetrics GitHub repository — Apache-2.0, created 2018-09-30; LTS backports of the v1.153.0 security fixes: v1.148.5, v1.136.19
- Single-node VictoriaMetrics documentation — flags, retention,
-search.latencyOffset, import paths, security recommendations (the basis for every default-value claim in this review, verified against-helpoutput on the live binary) - Cluster documentation — vmstorage/vminsert/vmselect architecture, replication, multitenancy, capacity planning
- vmagent documentation — per-URL persistent queues, FIFO replay,
-promscrape.configcompatibility, stream parsing - vmauth documentation — new-in-1.153.0 OIDC SSO, access control, URL routing
- MetricsQL documentation — no-extrapolation rate semantics,
keep_metric_names - VictoriaMetrics Enterprise features — downsampling, multiple retentions, automatic vmstorage discovery; pricing via the plans page, managed option at VictoriaMetrics Cloud
- Prometheus storage documentation — the local-TSDB durability and scaling limits that motivate this migration class
- vmctl migration tool and vmbackup — the moving-parts checklist for a real migration
- VictoriaMetrics case studies — vendor-published scale claims (Adidas, CERN, Brandwatch and others); treat as marketing until you reproduce them