OPA 1.21: The Indexer Rewrite, Two Silent Breaking Changes, and a Panic Patched in Five Days
Sources
- OPA v1.21.0 release notes
- OPA v1.21.1 release notes — comprehension regression hotfix
- Issue #5754 — YAML on/off parsed as booleans (YAML 1.1 trap)
- Issue #7275 — empty-literal type-checking gap
- Issue #9280 — the 1.21.0 compiler regression (panic + unsafe var)
- PR #9211 — rule_labels on Data and Query APIs
- PR #9205 — gzip bounded via klauspost gzhttp
- Issue #555 — stack traces for eval errors (opened 2018, closed by 1.21)
- Issue #9184 — config reload on --watch (still open after revert #9219)
- OPA docs — policy performance and rule indexing
Open Policy Agent 1.21 is a release about the tax you pay for policy breadth — and about two silent behavior changes that will bite anyone who upgrades without reading. We downloaded the actual v1.20.0, v1.21.0, and v1.21.1 binaries (sha256-verified against the release assets), ran the same probes against all three, and measured what the headline claim — a rewritten rule index — actually does to a policy that looks like production code. The honest summary: the indexer work is real and large, the YAML 1.2 core schema switch is a breaking change with no compile-time warning, the empty-literal type checker breaks previously-compiling policies, and the release shipped with a compiler bug so bad it panicked on valid Rego — patched five days later in v1.21.1. If you take one thing from this analysis: read the YAML section before you upgrade, and do not run 1.21.0 anywhere — start at 1.21.1.
The Pain: Linear Scans Over Policies That Grew Like Kudzu
Rego policies at platform scale do not look like the docs' examples. A Kubernetes admission policy library like Gatekeeper constraint templates, or a Terraform plan policy (the same OPA-in-CI pattern we mapped in our OpenTofu vs. Terraform comparison), accumulates hundreds or thousands of rule bodies under a single decision path, most of them guarded by string prefix tests. Before 1.21, OPA's rule index only indexed a narrow set of shapes: ground == comparisons and a couple of others. Everything else — every startswith guard, every reference rooted at a local variable, every lookup into base data — meant the evaluator fell back to linear evaluation over all candidate bodies on every query. On a warm server taking thousands of authz calls a second, that is the difference between an index lookup and a scan.
That is the industry pain the 1.21 index rewrite targets, and it is why the changelog's dry line — "improved rule indexing" — undersells the work. The rewrite makes the index handle shapes it previously could not:
startswithandendswith, plusstrings.any_prefix_match/strings.any_suffix_match, are now indexable when the base strings are known at compile time (PR #9161, #9164).- References rooted at a local variable (
x := input; x.foo == "a") now index like the direct form (PR #9081). - A lookup into base data (
data.groups.admins.members[input.subject]) indexes by asking the object for the key (PR #9235) — such rulesets used to leave every body a candidate. - The trie is smaller (unconstrained refs kept out of it, paths stop at the last constrained level), and candidate lookup collects candidates in a bitset (PR #9190, #9244).
query: data.perf.allow (input.path = "/other/thing")
┌──────────────────────────── OPA v1.20.0 ────────────────────────────┐
│ rule index: only ground == shapes indexed │
│ startswith(input.path, "/svc0000") ── not indexable ──┐ │
│ startswith(input.path, "/svc0001") ── not indexable ──┤ │
│ ... ├─▶ all 5,001
│ startswith(input.path, "/svc4999") ── not indexable ──┘ │
│ input.path == "/nothing" ── indexed (1 body) │
│ │
│ result: linear scan → timer_rego_query_eval_ns = 13,200,083 │
└────────────────────────────────────────────────────────────────────┘
┌──────────────────────────── OPA v1.21.x ───────────────────────────┐
│ rewritten index: ground prefixes of startswith/endswith indexed │
│ trie: "/svc0000" ──▶ body #1 (ground prefix lookup) │
│ ... │
│ "/svc4999" ──▶ body #5001 │
│ "/other/thing" matches nothing in the trie → 0 candidates │
│ │
│ result: trie lookup → timer_rego_query_eval_ns = 19,834 │
└────────────────────────────────────────────────────────────────────┘The Fix, Measured: 13.2 ms → 0.02 ms on a 5,000-Body Policy
Here is our probe, exactly reproducible with the commands below. We generated a synthetic policy with 5,000 prefix-guarded bodies plus one ground-equality body, ran the same non-matching query through v1.20.0, v1.21.0, and v1.21.1 with --metrics, and confirmed the matching-input case still returns true (correctness preserved — the index didn't just start dropping rules):
# generate 5,001-body policy: one == body, 5,000 startswith bodies
python3 - <<'EOF'
n = 5000
out = ["package perf", "", "allow if { input.path == \"/nothing\" }"]
for i in range(n):
out.append("allow if { startswith(input.path, \"/svc%04d\") }" % i)
open("/tmp/perf.rego", "w").write("\n".join(out))
EOF
# same query, same input, three versions — non-matching input
echo '{"path":"/other/thing"}' > /tmp/no-match.json
opa eval --data /tmp/perf.rego --input /tmp/no-match.json --metrics --format pretty 'data.perf.allow'
# v1.20.0: timer_rego_query_eval_ns 13,200,083 (undefined)
# v1.21.0: timer_rego_query_eval_ns 19,834 (undefined)
# v1.21.1: timer_rego_query_eval_ns 25,913 (undefined)
# correctness check: matching input must still evaluate to true
echo '{"path":"/svc0420/called"}' > /tmp/match.json
opa eval --data /tmp/perf.rego --input /tmp/match.json --format pretty 'data.perf.allow'
# v1.20.0: true v1.21.1: true ← index change never altered the verdictThat is a ~660x reduction in query-eval time on this shape (13.2 ms → 0.02 ms). We call it a probe, not a benchmark: it is one synthetic shape on one host, and your policy mix will differ. But the direction and magnitude are unambiguous, and they line up with the project's own performance guidance, which now documents the wider set of indexable statements. Two honesty notes from our own numbers: policy load + compile dominates cold-start (78–98 ms in our runs, essentially unchanged across versions), so CLI one-shot latency doesn't improve — the win is per-query on a warm opa run --server or sidecar; and the win only applies to ground prefix tests. If your prefixes are computed at eval time, keep your expectations flat.
Breaking Change #1: YAML Is Now the 1.2 Core Schema — Silently
The changelog line is three words — "YAML is parsed against the 1.2 core schema" (issue #5754, 3.5 years old). The operational reality is a silent data-shape change everywhere OPA reads YAML: --data files, bundles, config files, and the yaml.unmarshal builtin. Under YAML 1.1 (go-yaml v2), the bare words yes, no, on, off, y, n resolved to booleans. Under the 1.2 core schema they are plain strings. We probed it:
input YAML: on: push
env:
enable: yes
disable: no
OPA v1.20.0: { "env": { "disable": false, "enable": true }, "true": "push" }
OPA v1.21.1: { "env": { "disable": "no", "enable": "yes" }, "on": "push" }Look at the v1.20.0 output again: the GitHub Actions on: key — the single most common key in every workflow file — was parsed as boolean true. Any policy matching input.workflow["on"] was silently matching nothing. That is the trap the 1.2 switch closes, and it's worth saying the maintainers chose the correct side of this: YAML 1.2 core schema is what every other modern tool in the CNCF ecosystem already implements. But the migration tax is yours:
- Data files with bare
yes/no/on/offvalues become strings — if any policy compares them totrue/false, those comparisons now silently fail (boolean ≠ string in Rego). - Config files — the same file that configures your server — are parsed with the new schema. A
label: onstanza becomes the string"on". - The release notes' own mitigation is the whole fix: "quote the value and use true/false instead." There is no version of this where you don't grep your YAML estate first. If your policies ingest GitHub Actions workflows, Kubernetes manifests with legacy
on/offkeys, or any YAML from teams that never got the memo about booleans, run your test suites against 1.21 before the upgrade, not after.
Breaking Change #2: Empty Literals Are Now Typed as Empty
The second silent break is a type-checker change (issue #7275): {}, [], and set() literals are now typed as empty collections instead of "could hold anything." Previously obj := {}; obj.bar compiled fine. Now:
package play
p if {
obj := {}
obj.bar # v1.20.0: compiles, undefined at eval
} # v1.21.x: rego_type_error: undefined ref: obj.bar
# have: "bar" want (one of): []
# fix pattern: test emptiness without asserting the type
q if {
count(obj) == 0 # still compiles on both
}We verified this on all three binaries: v1.20.0 compiles the left side; v1.21.0 and v1.21.1 both reject it with rego_type_error. The changelog is explicit that comparing a non-empty literal against {} also becomes a match error, and that some x in [] now fails to compile too. This is a compile-time break, so it surfaces in CI (opa check exits 1) rather than in production eval paths — if you gate your policies through opa check in CI, you'll catch it before deploy. If you don't, the first thing to break is your bundle build. Either way, the previously idiomatic "try the key, fall back" patterns die here.
The Recall: 1.21.0 Shipped a SIGSEGV, 1.21.1 Fixed It in Five Days
Now the part that should shape your upgrade decision more than anything else above. On September 29 — five days after 1.21.0 — the project shipped v1.21.1 with a single fix: a comprehension using some … in or every, nested inside an object or set literal, was wrongly treated as "ground," so the compiler skipped rewriting it. The result was rego_unsafe_var_error for the common form and a compiler panic — a genuine nil-pointer SIGSEGV — for the every form. We reproduced both on the v1.21.0 binary:
policy: p := {"k": [r.a | some r in input.xs]} # unsafe var
q := {[1 | every x in input.xs { x > 0 }]} # SIGSEGV
v1.21.0 output:
1 error occurred: comp.rego:3: rego_unsafe_var_error: var r is unsafe
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0xb38a57]
goroutine 1 [running]:
github.com/open-policy-agent/opa/v1/topdown.(*bindings).plugNamespaced(...)
/src/v1/topdown/bindings.go:94 +0x17Both cases evaluate correctly on v1.21.1 (and both worked on v1.20.0). Reported by @tun0 on September 29 and fixed by Stephan Renatus (@srenatus) the same day — that turnaround is the project working as intended, and it is exactly why 1.21.0 should be treated as a do-not-run artifact. If your bundle pipeline caches base images with OPA pinned by minor version, now is the moment to find out whether "1.21" in your Dockerfile actually resolved to 1.21.0 during the five-day window.
Hidden Costs and the Things the Changelog Buries
Every release has costs. This one has a few worth calling out separately:
- The
--watchconfig reload that wasn't. Theopa run --helptext in 1.21 claims it watches "the configuration file for changes" — we verified the text is there. But the underlying feature was reverted before release (#9219) (see the changelog entry "runtime: Revert config file reload on --watch"), and we confirmed empirically: editing the config'sgzip.min_lengthunder a running--watchserver processed watch events in the logs but never applied the new value. The issue tracking the real feature, #9184, remains open. If you scripted config hot-reload around the 1.21 announcement, unscript it. - Recursive rules with variable heads stay unsupported on wasm/IR targets. The recursion check got smarter (#6813 — dotted heads with variables no longer false-positive as recursive), but we verified the fix is interpreter-only:
opa build -t wasmstill errors on the same policy with "rules sharing that path prefix are planned as a single function." If you target wasm SDKs, don't write variable-head recursive rules yet. - The gzip rewrite is behavior-compatible but subtly different. The server's gzip path was rebuilt on klauspost/compress's gzhttp (PR #9205), capping the pre-compression buffer at
min_length(floored at 512 bytes) instead of buffering an entire Write call. We probed the floor: withmin_length: 10, a 16-byte response still compresses — the floor bounds the buffer, not the threshold. Only gzip is negotiated; zstd isn't. - Stack traces arrived for eval errors — eight years after the request. Issue #555 was opened January 11, 2018 and finally closed by this release; evaluation errors now carry structured stack traces, and the debugger gained a
querystack-trace framing mode (PR #9128). For anyone who has debugged a production policy incident with nothing but "undefined," this is the sleeper feature of the release. - OCI bundle downloads got fixes that matter at scale — image-index resolution behind an index (#7461, a year and a half old), a
Trigger()race that could report false success, and ignored OCI downloader settings (#9113). - Decision-log masking got a data-race fix and requeued-chunk upload failures are now reported (#9186) — small lines, but they're the difference between "logs are complete" being a true statement and an aspiration.
API Additions You Can Build On: rule_labels
One quiet API addition will matter to anyone building audit tooling: the Data API (/v1/data) and Query API (/v1/query) now accept a rule_labels query parameter that returns the merged # METADATA labels of the rules that fired, in the response payload (PR #9211). Previously those labels existed only in decision-log events. We probed it against a live server:
# policy with metadata labels:
# METADATA
# labels:
# severity: high
allow if input.user == "alice"
# v1.20.0 server:
$ curl -s -X POST -H 'Content-Type: application/json' \
-d '{"input":{"user":"alice"}}' \
'http://localhost:8181/v1/data/play/allow?rule_labels=true'
{"result":true} # ← parameter ignored
# v1.21.1 server, same policy, same request:
{"result":true,"rule_labels":[{"severity":"high"}]}If your compliance story is "we know which rule allowed it," you no longer need to ship and parse decision logs to answer the follow-up question — what did that rule claim about itself? The label schema is yours to define; use it.
Upgrade Runbook: Find Your YAML and Your 1.21.0s
This is a minor release, so most fleets will take it via their distribution (Gatekeeper bundles its own OPA — check the Gatekeeper release notes for its cadence) or their base image. For teams running standalone OPA, the binary swap is trivial; the migration is the YAML estate and the pin audit:
# 1. Find YAML that relies on 1.1 booleans (data dirs, config, values):
grep -rInE ':\s*(yes|no|on|off|y|n)\s*$' policies/ config/ data/
# → every hit is a silent type change; quote it or convert to true/false
# 2. Compile-check every policy against the new type checker (1.21.1):
opa check --strict policies/ # exits 1 on the new rego_type_errors
# 3. Audit for the five-day 1.21.0 window (Sep 24-29, 2026):
# grep base-image pins across Dockerfiles / CI / helm values
grep -rn --include='Dockerfile*' --include='*.yaml' --include='*.yml' \
-E '(openpolicyagent/opa|.*opa[^[:alnum:]])[:@v]?1\.21\.0' .
# anything that resolved 'opa:1.21' to 1.21.0 during the window needs a
# rebuild on 1.21.1 — do not trust cached layer digests
# 4. Binary upgrade (standalone server), sha256-verified:
curl -fLO https://github.com/open-policy-agent/opa/releases/download/v1.21.1/opa_linux_amd64
curl -fLO https://github.com/open-policy-agent/opa/releases/download/v1.21.1/opa_linux_amd64.sha256
sha256sum -c opa_linux_amd64.sha256 # 23c654cd…f48fbb
install -m 0755 opa_linux_amd64 /usr/local/bin/opa
opa versionOne more note for CI: if you lint Rego, the Regal linter is the community-standard companion tool, and its nightly compatibility run against OPA main is now part of OPA's own CI — a good sign for catching the next 1.21.0-class regression before it ships.
Verdict
Upgrade — to 1.21.1, after a YAML audit, not before. The indexer rewrite is the best performance news OPA has shipped in years: a ~660x eval-time reduction on a shape that production policies actually have, with correctness intact and the project's own docs updated to match. The two breaking changes are both defensible engineering — YAML 1.2 is where the ecosystem already was, and empty-literal typing makes the compiler catch what it previously let through — but neither warns you at upgrade time; the YAML one doesn't even fail a compile, it just quietly changes what yes means. Run opa check against every policy and grep every YAML file that touches OPA before you ship this. And treat 1.21.0 as radioactive: the panic was real, we reproduced it, and the fix shipped in five days — which is both a compliment to the maintainers and a reminder that a .0 release from any project is a bet, not a version.
Credit Where Due
The index rewrite is overwhelmingly the work of Stephan Renatus and Stephen Spaink — #9161, #9164, #9235, #9081, #9190 and the compiler-side type-checker work respectively — with Tobias Stockert on the prefix-indexing PR and Anders Eknert across the AST cleanups and the Regal compatibility work. Johan Fylling shipped the debug stack-trace framing, Charlie Egan modernized the website (local Pagefind search, AI-guidelines updates), and the 1.21.1 hotfix came from the same Stephan Renatus, from a report by @tun0. Community reporters who moved this release: @scnewma, @johnc-c (the YAML trap), @disaverio (#7275), @vlsi, @zscott (#7461), @charlesdaniels, @kmadan, @rothenes, @kishorviswanathan (#6854), @nikpivkin, @stobias123, @msahmi. The upstream issues are 3.5 years old in the YAML case — this is what maintenance on a graduated CNCF project actually looks like.