Grafana Cloud vs. Datadog: Cost, Lock-in, and the Self-Hosted Trade-off

Sources

You are about to sign a multi-year observability contract and the shortlist is down to two names: Grafana Cloud and Datadog. This guide answers the procurement question directly — what each actually costs at your scale, how locked in you get, and when the third option (self-hosting the Grafana stack) beats both. Every price below was verified from the two vendors' own pricing pages on 2026-09-28 — treat any number older than a quarter as stale and re-verify before signing. We have no financial relationship with either vendor; all links are direct and untracked, and neither company saw this page before publication.

The uncomfortable finding up front: the folklore on both sides is wrong. Datadog is not automatically the expensive one, and Grafana Cloud is not automatically the cheap one. At list prices, a 100-node Kubernetes stack with 400k active series and 300 GB/day of logs costs ~$5,270/month on Datadog and ~$5,750/month on Grafana Cloud under identical assumptions — Grafana is ~9% dearer at list, until you pull its built-in cost levers, at which point it lands ~30% under Datadog. The vendors bill fundamentally different things, so which one wins depends on what your telemetry looks like, not on a sticker price. The worked math is below; copy it and rerun it with your own numbers.

TL;DR — the 30-second verdict

Your situation Pick Why
You want one vendor for everything — infra, APM, logs, RUM, CI visibility, security — and you have the budget for it Datadog The broadest integrated product surface in the market; sixth consecutive Gartner MQ Leader (July 2026). You pay per host and you pay for breadth
You already run Prometheus and Loki, or your metrics are high-cardinality and your logs are high-volume Grafana Cloud Same protocols and APIs as the OSS stack; usage-based billing that rewards tuning (Adaptive Metrics/Logs); the AGPL escape hatch keeps exit realistic
You have platform engineers with spare capacity and hard cost pressure at serious scale Self-hosted Grafana stack Loki + Mimir + Tempo + Grafana are all AGPL-3.0 and free; a single well-sized node beats both vendors' list prices below ~150k series — but budget an FTE, not a server
You are a small team (<10 hosts) evaluating for the first time Grafana Cloud Free/Pro Free tier carries 10k series and 50 GB/mo of each signal with no time limit; Pro is $19/mo platform fee plus usage. Datadog's free tier holds 5 hosts for 1-day retention
Your compliance posture requires vendor SOC 2 / FedRAMP-style deployment options Both — push to Enterprise tier Grafana offers Federal Cloud and Bring-Your-Own-Cloud deployment at its $25k/yr Enterprise tier; Datadog has established compliance programs. Neither is a differentiator at the low tiers

The two billing models, translated

Every bill shock story about these two vendors traces back to one sentence neither pricing page says plainly: they meter different scarce resources. Datadog bills hosts and events — countable things you can divide by your fleet size. Grafana Cloud bills usage — series, gigabytes, and host-hours, which vary with how much telemetry each unit emits. Neither model is honest about its failure mode: Datadog's punishes fleets of many small hosts, and Grafana's punishes cardinality spikes and chatty logs.

Signal Datadog bills you for (annual list, US) Grafana Cloud bills you for (Pro list)
Infrastructure metrics $15/host/mo Pro, $23 Enterprise (on-demand: $18/$27). 100–200 custom metrics per host included; beyond that, $5 per 100/mo. 15-month retention included $6.50 per 1k billable series (first 10k free). Billable series = max(active series, total data points per minute) — so scrape interval matters as much as series count. 13 months retention included
Containers 5 containers per Pro host license included, then $1/container/mo (or $0.002/hr on-demand) No container license — Kubernetes Monitoring bills $0.01/host-hour ($7.20/host/mo) plus $0.0007/container-hour, first 2,232 host-hours free
Logs Ingest $0.10/GB plus indexed events: $1.70 per 1M events at 15-day retention (on-demand $2.55; 3-day $1.06, 7-day $1.27, 30-day $2.50). Flex Logs: $0.05 per 1M events stored/mo for cold tiers Three-part meter: Process $0.050/GB + Write $0.400/GB (what lands in the database) + Retain $0.100/GB (storage accrual). First 50 GB/mo free, 30-day retention on Pro
Traces / APM $31/APM host/mo (with infra attached; $36 on-demand), APM Pro $35, APM Enterprise $40. Standalone APM without infra: $36–47. OTLP ingest: $0.10/GB. Indexed spans beyond the included 1M/host/mo: $1.70 per 1M at 15-day $0.050/GB processed (traces share the logs-style process meter), first 50 GB/mo free. Or App Observability at $0.025/host-hour ($18/host/mo), first 2,232 host-hours free
Platform floor None — per-host pricing starts at 1 host. Free tier: 5 hosts, 1-day retention $19/mo Pro platform fee. Free tier: 10k series, 50 GB/mo per signal, 14-day retention, no time limit
Enterprise floor Volume discounts start around 500+ hosts/mo; custom contracts $25,000/yr minimum spend commit — that is the gate for premium support, custom retention, Federal/BYOC deployment

Two structural details that routinely surprise buyers, both from vendor documentation: Grafana's billable series is not your active series count. It is max(active_series, total_DPM) — if you scrape at 15-second intervals (4 DPM), a 100k-series job bills like 400k series. The docs also confirm billing uses the 95th percentile of that maximum, forgiving roughly the top 36 hours of each month — which cuts both ways: a spike under 36 h is forgiven, but a spike over 36 h sets your whole month. Datadog's container math bites Kubernetes fleets specifically: the 5-containers-per-host allotment predates the era of 30-pod nodes, so a dense cluster pays $1/container/mo for most of its pods, quietly adding hundreds of dollars per node per year on top of the host license.

What the same stack actually costs — worked example

Assumptions, stated plainly so you can rerun this with your own numbers: 100-node Kubernetes cluster, 60 of those nodes running traced services, ~10 extra billable containers per node beyond allotments, 300 GB/day (9 TB/mo) uncompressed log ingest, 30 GB/day of tail-sampled traces, 400k active Prometheus series at 1 DPM, 20% of Datadog log volume indexed at 15-day retention, 50% of Grafana log volume written after Adaptive Logs, all list prices, annual billing. Every input is a lever — the point is the method, not the total.

Line item Datadog (annual list) Grafana Cloud (Pro list)
Infrastructure / metrics base $1,500 — 100 × $15 Infra Pro $2,535 — 390k billable series × $6.50/1k
APM / traces $1,860 — 60 × $31 APM (infra attached) $495 — Tempo process/write/retain on 900 GB/mo
Containers $1,000 — 10 extra × 100 hosts × $1 — (no container license on the standalone path)
Logs $909 — $900 ingest + $9 index (5.4M events @15d) $2,700 — $450 process + $1,800 write (50%) + $450 retain
Platform fee — $19
Monthly total $5,269 $5,749
Annualized ~$63,230 ~$68,990

Worked example, list prices verified 2026-09-28 — not a quote. Volume discounts (Datadog at 500+ hosts; Grafana's automatic tiered discounts) and committed-spend pricing change both columns materially at real scale. Run your own numbers before you believe any vendor's comparison page — including Grafana's own Grafana vs. Datadog page, which, like everything on a vendor's /compare, is marketing.

Read that table twice before you forward it, because the result inverts at both ends. The same inputs on a 10-node cluster (6 traced hosts, 30 GB/day logs, 60k series) give $527/month on Datadog vs. $664/month on Grafana Cloud — Datadog's floor is simply lower, and Grafana's free-tier allowances are too small to close the gap. But shift the inputs toward high cardinality and heavy logs — the profile a platform team actually runs — and the meters diverge hard. The bill-shock risk lives in different places for each vendor. Datadog's is the custom-metric meter: 400k custom metrics on a 100-host plan (200/host allotted) bill $19,000/month on the classic meter — the single most expensive line in this entire comparison, and the reason the "Datadog is expensive" folklore exists (the $0.10/100 ingested-custom-metrics meter cuts that hazard to $400, but you must know to switch). Grafana's is the DPM trap: billable series = max(active series, total data points per minute), so scraping 400k series at 15-second intervals (4 DPM) bills like 1.6M series — $10,400/month on the metrics line alone. And because billing uses the 95th percentile of the month, a cardinality spike that outlasts ~36 hours sets your whole month.

The real differentiator is the cost-control tooling each vendor ships, because both let you cut the bill after the fact — differently. Grafana builds the levers into the platform: Adaptive Metrics (vendor claims up to 80% metrics reduction) aggregates away unused series; Adaptive Logs (claims up to 50% ingest reduction) drops and samples noisy patterns. Applied to the worked example, cutting 400k active series to 80k with Adaptive Metrics saves $2,080/month — dropping Grafana Cloud from $5,749 to $3,669, now 30% under Datadog for the same stack. Datadog's equivalents (ingest controls, index-rate limiting, the ingested-custom-metrics meter) are effective but manual: you write the configuration, you own the filters, and you field the on-call complaint when something someone wanted lands in the "not indexed" bucket. That is the honest summary of the operational trade: Grafana automates cost optimization and bakes it into the platform; Datadog hands you the knife and expects you to cut.

Lock-in: AGPL in, Apache out — and why it matters for exit

Lock-in is where these two are genuinely different, and it is not about export buttons. Every core Grafana Labs backend — Loki (logs), Mimir (metrics), Tempo (traces), Pyroscope (profiles), and the Grafana dashboard server itself — is AGPL-3.0 (verified via the GitHub API, 2026-09-28). Datadog's agent is Apache-2.0, but it is the only open component: the entire query, storage, dashboarding, and correlation layer that makes Datadog worth its price is closed SaaS. That asymmetry defines your exit story.

If you choose Grafana Cloud and later want out, you are not leaving a product — you are leaving a hosting arrangement. The self-hosted stack speaks the same APIs (Prometheus remote_write, the Loki push API, OTLP), stores data in your own object storage in the same formats, and runs the same dashboards. Grafana Labs is a hosting company that open-sources its moat; the moat still exists (managed scaling, Adaptive features, IRM, support), but you can walk. If you choose Datadog, the data leaves with you only as exports and archives (S3 rehydration is well supported), and every dashboard, monitor, and team workflow you built is gone — the cost of leaving is not the export, it is rebuilding years of operational muscle memory in a new tool.

Lock-in dimension Grafana Cloud Datadog
Core backend license AGPL-3.0 (Loki, Mimir, Tempo, Pyroscope, Grafana) — all runnable by you, forever Agent Apache-2.0; everything else closed SaaS
Data export Open formats (Prometheus TSDB, Loki chunks in your object storage, OTLP in) APIs + S3 archives; rehydration supported, but re-import to a competitor is a project
Dashboards as code JSON dashboards provisioned via Git — portable to any Grafana, self-hosted included Dashboard JSON exists, but the runtime is proprietary; port means rewrite
Instrumentation portability OpenTelemetry first-class; Prometheus-native. Switching vendors = repointing remote_write OTLP supported, but the differentiating features (Watchdog, correlations, USM) are proprietary analysis of your data
Realistic exit cost Low–medium: rehost the same software, migrate storage, keep dashboards High: tooling rebuild, retraining, historical data stays behind

One nuance the AGPL story omits: AGPL is not a free-software halo, it is a network-copyleft obligation. If you self-host and modify Loki or Grafana and expose the service over a network — including inside your company — you must offer the modified source to users of that service. For internal platform teams this is usually a non-issue, but it is a real lawyer-conversation if you embed Grafana into a customer-facing product. Datadog's closed stack has no such clause because there is nothing to modify. Neither vendor's license constrains your telemetry pipeline; both accept OpenTelemetry, which is the actual portability layer worth standardizing on regardless of which one you pick.

The third option: self-hosting the Grafana stack

Every observability procurement eventually asks the CFO question: why are we paying a vendor at all when the underlying software is free? The honest answer has three parts. First, the software really is free — Loki, Mimir, Tempo, Pyroscope, and Grafana are all AGPL-3.0 with no open-core feature gating on the storage path. Second, the community datapoint is real but dated: a widely-cited HN report from 2022 ran Grafana + Prometheus with 141k series on a single t3.xlarge (~$100/month all-in). Third — and this is the part vendor comparison pages never include — the bill you stop paying the vendor, you start paying in staff time. A self-hosted observability stack at mid-size is a 0.5–1 FTE commitment across upgrades, storage capacity planning, HA, and the 3 a.m. "Loki is down and Loki has the logs" failure mode (we measured the storage side of that trade in Loki vs. Elasticsearch on real disk). Self-hosting wins on cost when your telemetry volume is high and stable, your team already runs Kubernetes well, and you can size the operational tax honestly; it loses when volume is spiky or the team is small.

For the worked example's scale (400k series, 9 TB/mo logs), self-hosting on cloud VMs and object storage lands in the $800–1,500/month infrastructure range — roughly a quarter of either vendor's $5,300–5,700 — before the staffing line. That gap is the entire open-source observability business model: both vendors are, at root, selling you the operational team you did not want to hire. The decision is not "cloud vs. self-hosted" as an ideology; it is whether your platform team's marginal hour is cheaper than the vendor's margin.

# Same agent fleet, same dashboards, same PromQL — the migration is a URL.
# Grafana Cloud host (what you leave):
remote_write:
  - url: https://prometheus-prod-01-eu-west-0.grafana.net/api/prom/push
    basic_auth:
      username: <instance-id>
      password_file: /etc/secrets/grafana_cloud_api_key

# Self-hosted Mimir (what you land on — AGPL-3.0, your object storage):
remote_write:
  - url: http://mimir-gateway.observability.svc:80/api/v1/push
    # storage: your S3/GCS bucket, your retention policy, your cost curve
    queue_config:
      max_samples_per_send: 5000
      capacity: 10000

Verdicts — who should sign with whom

Datadog — the all-in platform, priced per host

What it is: the broadest integrated observability platform in the market — infrastructure, APM, logs, RUM, synthetics, CI visibility, security (CSM/Cloud SIEM), and increasingly AI-agent monitoring, all behind one agent and one query surface. Named a Leader in the Gartner MQ for Observability Platforms for the sixth consecutive year (July 2026).

Honest strength: the correlation experience. Pivoting from a slow trace to the exact container's logs, processes, network flows, and the deploy that shipped it is genuinely faster than stitching that story together in Grafana across four data sources. Watchdog's anomaly detection is real product, not marketing. For a 20–200 host org without a dedicated observability team, Datadog replaces an entire subsystem of internal tooling.

Honest weakness: the meters. Custom metrics beyond the 200/host allotment at $5 per 100/month, containers beyond 5–10/host at $1 each, indexed log events at $1.70 per 1M at 15-day — every one is a meter that quietly scales with the messiness of your stack, and the community consensus (and years of bill-shock threads) is that Datadog makes it easy to misconfigure and expensive to find out. Procurement without a usage cap in the contract is how the folklore got written.

Pricing as of 2026-09-28 (datadoghq.com/pricing, full unit list — verify with vendor): Infra Pro $15/host/mo, Enterprise $23; APM $31–40/host; Log ingest $0.10/GB + index $1.06–2.50 per 1M events by retention; free tier capped at 5 hosts with 1-day metric retention; 14-day full-product trial.

Who should pick it: teams that want one throat to choke, value breadth over price, run moderate-cardinality workloads, and will negotiate a committed-spend contract with caps before signing. Who should skip it: high-cardinality Prometheus shops (the custom-metrics meter will find you), cost-sensitive platform teams already fluent in OSS observability, and anyone who needs on-prem — there is no self-hosted Datadog.

Grafana Cloud — the usage-metered host of an open stack

What it is: the managed version of the open-source stack you may already run — Prometheus/Mimir metrics, Loki logs, Tempo traces, Pyroscope profiles, Grafana dashboards — plus managed extras (Adaptive Metrics/Logs, Kubernetes and app observability bundles, IRM, Grafana Assistant and Investigations on the AI side). Also a Gartner MQ Leader in 2026. Product and docs: grafana.com/products/cloud.

Honest strength: the exit is real. Everything underneath is AGPL-3.0 and speaks native Prometheus and OTLP — adopting Grafana Cloud is a hosting decision, not a platform bet, and the pricing model (usage with automatic volume discounts, generous free tier, and built-in cost-reduction tooling) rewards the tuning work a good platform team does anyway. Adaptive Metrics alone can halve a metrics bill in an afternoon.

Honest weakness: the meters are different, not absent. Billable series = max(active series, DPM), billed at the 95th percentile of the month — a scrape-interval mistake or a cardinality spike that outlasts 36 hours sets your whole month, and the three-part log meter (process/write/retain) means the same GB can cost $0.05 or $0.55 depending on retention and filtering. Integration is also more assembly than product: correlating a trace to the exact log line to the profile is possible but you assemble it, where Datadog hands it to you pre-wired. Community complaints about series-based pricing exist here too — the grass is not categorically greener.

Pricing as of 2026-09-28 (grafana.com/pricing — verify with vendor): Free tier (10k series, 50 GB/mo per signal, 14-day retention, no time limit); Pro $19/mo platform fee + usage (metrics $6.50/1k series, logs $0.050 process + $0.400 write + $0.100 retain per GB, traces $0.050/GB); Enterprise from $25,000/yr commit (premium support, custom retention, Federal Cloud/BYOC).

Who should pick it: Prometheus/Loki-native teams, high-log-volume or high-cardinality shops that will actually use Adaptive Metrics/Logs, orgs with a plausible self-hosting future, and cost-sensitive buyers below Enterprise scale. Who should skip it: teams that want a fully pre-wired experience with zero assembly, orgs with no one to own dashboards-as-code, and anyone for whom the $25k/yr Enterprise floor is the only acceptable tier (at that level, negotiate both vendors).

Self-hosted Grafana stack — the engineer's option

What it is: Loki + Mimir + Tempo + Pyroscope + Grafana on your own Kubernetes, storage in your own object store, all AGPL-3.0.

Honest strength: the cost curve. Below ~150k series and a few TB of logs, a single well-sized node beats both vendors' list prices by an order of magnitude, and there is no meter watching your cardinality. The 2022 community datapoint (141k series, one t3.xlarge, ~$100/mo) remains directionally true at that scale.

Honest weakness: you are the vendor now. High-availability Mimir and Loki are genuinely hard to operate well — multi-tenancy, compactor/distributor tuning, object-storage lifecycle, and upgrade discipline are a standing 0.5–1 FTE load, and "Loki is down and Loki has the logs" is a real outage mode you get to own. AGPL also means network-copyleft obligations if you modify and share.

Pricing: $0 license; infrastructure cost only — realistically $800–1,500/mo of cloud resources at the worked-example scale, plus the staffing line that dwarfs it.

Who should pick it: platform teams with Kubernetes competence, stable high volume, and a compliance need to own the data plane. Who should skip it: everyone else, especially teams under 20 hosts — the operational tax exceeds the vendor premium at small scale.

Bottom line

If you are a platform team already running Prometheus and Loki, pick Grafana Cloud — the pricing model rewards the tuning you already do, the free tier lets you prove it on real traffic before paying, and the AGPL exit keeps your negotiating position honest. If you are a product-led org that wants observability to be someone else's problem and can cap the meters in the contract, pick Datadog — the integrated experience is the best on the market and the premium is the price of not building an observability team. If you have the staff and the scale, self-host the AGPL stack and pay yourself instead of a vendor — but budget the FTE honestly, because the failure mode of self-hosted observability is never the software; it is the team that no longer has time for it.

Whatever you sign, standardize the pipeline on OpenTelemetry and Prometheus protocols before you standardize the vendor. The telemetry layer is the part you can own permanently; the vendor is the part you will renegotiate.