OPA 1.21: The Indexer Rewrite, Two Silent Breaking Changes, and a Panic Patched in Five Days

Sources

Open Policy Agent 1.21 is a release about the tax you pay for policy breadth — and about two silent behavior changes that will bite anyone who upgrades without reading. We downloaded the actual v1.20.0, v1.21.0, and v1.21.1 binaries (sha256-verified against the release assets), ran the same probes against all three, and measured what the headline claim — a rewritten rule index — actually does to a policy that looks like production code. The honest summary: the indexer work is real and large, the YAML 1.2 core schema switch is a breaking change with no compile-time warning, the empty-literal type checker breaks previously-compiling policies, and the release shipped with a compiler bug so bad it panicked on valid Rego — patched five days later in v1.21.1. If you take one thing from this analysis: read the YAML section before you upgrade, and do not run 1.21.0 anywhere — start at 1.21.1.

The Pain: Linear Scans Over Policies That Grew Like Kudzu

Rego policies at platform scale do not look like the docs' examples. A Kubernetes admission policy library like Gatekeeper constraint templates, or a Terraform plan policy (the same OPA-in-CI pattern we mapped in our OpenTofu vs. Terraform comparison), accumulates hundreds or thousands of rule bodies under a single decision path, most of them guarded by string prefix tests. Before 1.21, OPA's rule index only indexed a narrow set of shapes: ground == comparisons and a couple of others. Everything else — every startswith guard, every reference rooted at a local variable, every lookup into base data — meant the evaluator fell back to linear evaluation over all candidate bodies on every query. On a warm server taking thousands of authz calls a second, that is the difference between an index lookup and a scan.

That is the industry pain the 1.21 index rewrite targets, and it is why the changelog's dry line — "improved rule indexing" — undersells the work. The rewrite makes the index handle shapes it previously could not:

            query: data.perf.allow  (input.path = "/other/thing")

  ┌──────────────────────────── OPA v1.20.0 ────────────────────────────┐
  │  rule index: only ground == shapes indexed                          │
  │     startswith(input.path, "/svc0000")   ── not indexable ──┐       │
  │     startswith(input.path, "/svc0001")   ── not indexable ──┤       │
  │     ...                                                          ├─▶ all 5,001
  │     startswith(input.path, "/svc4999")   ── not indexable ──┘     │
  │     input.path == "/nothing"             ── indexed (1 body)       │
  │                                                                    │
  │  result: linear scan → timer_rego_query_eval_ns = 13,200,083      │
  └────────────────────────────────────────────────────────────────────┘

  ┌──────────────────────────── OPA v1.21.x ───────────────────────────┐
  │  rewritten index: ground prefixes of startswith/endswith indexed   │
  │     trie:  "/svc0000" ──▶ body #1      (ground prefix lookup)      │
  │             ...                                                    │
  │            "/svc4999" ──▶ body #5001                                │
  │     "/other/thing" matches nothing in the trie → 0 candidates      │
  │                                                                    │
  │  result: trie lookup → timer_rego_query_eval_ns = 19,834           │
  └────────────────────────────────────────────────────────────────────┘

The Fix, Measured: 13.2 ms → 0.02 ms on a 5,000-Body Policy

Here is our probe, exactly reproducible with the commands below. We generated a synthetic policy with 5,000 prefix-guarded bodies plus one ground-equality body, ran the same non-matching query through v1.20.0, v1.21.0, and v1.21.1 with --metrics, and confirmed the matching-input case still returns true (correctness preserved — the index didn't just start dropping rules):

# generate 5,001-body policy: one == body, 5,000 startswith bodies
python3 - <<'EOF'
n = 5000
out = ["package perf", "", "allow if { input.path == \"/nothing\" }"]
for i in range(n):
    out.append("allow if { startswith(input.path, \"/svc%04d\") }" % i)
open("/tmp/perf.rego", "w").write("\n".join(out))
EOF

# same query, same input, three versions — non-matching input
echo '{"path":"/other/thing"}' > /tmp/no-match.json

opa eval --data /tmp/perf.rego --input /tmp/no-match.json --metrics --format pretty 'data.perf.allow'
# v1.20.0: timer_rego_query_eval_ns  13,200,083   (undefined)
# v1.21.0: timer_rego_query_eval_ns      19,834   (undefined)
# v1.21.1: timer_rego_query_eval_ns      25,913   (undefined)

# correctness check: matching input must still evaluate to true
echo '{"path":"/svc0420/called"}' > /tmp/match.json
opa eval --data /tmp/perf.rego --input /tmp/match.json --format pretty 'data.perf.allow'
# v1.20.0: true    v1.21.1: true   ← index change never altered the verdict

That is a ~660x reduction in query-eval time on this shape (13.2 ms → 0.02 ms). We call it a probe, not a benchmark: it is one synthetic shape on one host, and your policy mix will differ. But the direction and magnitude are unambiguous, and they line up with the project's own performance guidance, which now documents the wider set of indexable statements. Two honesty notes from our own numbers: policy load + compile dominates cold-start (78–98 ms in our runs, essentially unchanged across versions), so CLI one-shot latency doesn't improve — the win is per-query on a warm opa run --server or sidecar; and the win only applies to ground prefix tests. If your prefixes are computed at eval time, keep your expectations flat.

Breaking Change #1: YAML Is Now the 1.2 Core Schema — Silently

The changelog line is three words — "YAML is parsed against the 1.2 core schema" (issue #5754, 3.5 years old). The operational reality is a silent data-shape change everywhere OPA reads YAML: --data files, bundles, config files, and the yaml.unmarshal builtin. Under YAML 1.1 (go-yaml v2), the bare words yes, no, on, off, y, n resolved to booleans. Under the 1.2 core schema they are plain strings. We probed it:

input YAML:      on: push
                 env:
                   enable: yes
                   disable: no

OPA v1.20.0:     { "env": { "disable": false, "enable": true }, "true": "push" }
OPA v1.21.1:     { "env": { "disable": "no",  "enable": "yes" }, "on": "push" }

Look at the v1.20.0 output again: the GitHub Actions on: key — the single most common key in every workflow file — was parsed as boolean true. Any policy matching input.workflow["on"] was silently matching nothing. That is the trap the 1.2 switch closes, and it's worth saying the maintainers chose the correct side of this: YAML 1.2 core schema is what every other modern tool in the CNCF ecosystem already implements. But the migration tax is yours:

Breaking Change #2: Empty Literals Are Now Typed as Empty

The second silent break is a type-checker change (issue #7275): {}, [], and set() literals are now typed as empty collections instead of "could hold anything." Previously obj := {}; obj.bar compiled fine. Now:

package play

p if {
    obj := {}
    obj.bar      # v1.20.0: compiles, undefined at eval
}                # v1.21.x: rego_type_error: undefined ref: obj.bar
                 #          have: "bar"   want (one of): []

# fix pattern: test emptiness without asserting the type
q if {
    count(obj) == 0     # still compiles on both
}

We verified this on all three binaries: v1.20.0 compiles the left side; v1.21.0 and v1.21.1 both reject it with rego_type_error. The changelog is explicit that comparing a non-empty literal against {} also becomes a match error, and that some x in [] now fails to compile too. This is a compile-time break, so it surfaces in CI (opa check exits 1) rather than in production eval paths — if you gate your policies through opa check in CI, you'll catch it before deploy. If you don't, the first thing to break is your bundle build. Either way, the previously idiomatic "try the key, fall back" patterns die here.

The Recall: 1.21.0 Shipped a SIGSEGV, 1.21.1 Fixed It in Five Days

Now the part that should shape your upgrade decision more than anything else above. On September 29 — five days after 1.21.0 — the project shipped v1.21.1 with a single fix: a comprehension using some … in or every, nested inside an object or set literal, was wrongly treated as "ground," so the compiler skipped rewriting it. The result was rego_unsafe_var_error for the common form and a compiler panic — a genuine nil-pointer SIGSEGV — for the every form. We reproduced both on the v1.21.0 binary:

policy:  p := {"k": [r.a | some r in input.xs]}    # unsafe var
         q := {[1 | every x in input.xs { x > 0 }]}  # SIGSEGV

v1.21.0 output:
  1 error occurred: comp.rego:3: rego_unsafe_var_error: var r is unsafe

  panic: runtime error: invalid memory address or nil pointer dereference
  [signal SIGSEGV: segmentation violation code=0x1 addr=0x0 pc=0xb38a57]

  goroutine 1 [running]:
  github.com/open-policy-agent/opa/v1/topdown.(*bindings).plugNamespaced(...)
          /src/v1/topdown/bindings.go:94 +0x17

Both cases evaluate correctly on v1.21.1 (and both worked on v1.20.0). Reported by @tun0 on September 29 and fixed by Stephan Renatus (@srenatus) the same day — that turnaround is the project working as intended, and it is exactly why 1.21.0 should be treated as a do-not-run artifact. If your bundle pipeline caches base images with OPA pinned by minor version, now is the moment to find out whether "1.21" in your Dockerfile actually resolved to 1.21.0 during the five-day window.

Hidden Costs and the Things the Changelog Buries

Every release has costs. This one has a few worth calling out separately:

API Additions You Can Build On: rule_labels

One quiet API addition will matter to anyone building audit tooling: the Data API (/v1/data) and Query API (/v1/query) now accept a rule_labels query parameter that returns the merged # METADATA labels of the rules that fired, in the response payload (PR #9211). Previously those labels existed only in decision-log events. We probed it against a live server:

# policy with metadata labels:
# METADATA
# labels:
#   severity: high
allow if input.user == "alice"

# v1.20.0 server:
$ curl -s -X POST -H 'Content-Type: application/json' \
    -d '{"input":{"user":"alice"}}' \
    'http://localhost:8181/v1/data/play/allow?rule_labels=true'
{"result":true}                          # ← parameter ignored

# v1.21.1 server, same policy, same request:
{"result":true,"rule_labels":[{"severity":"high"}]}

If your compliance story is "we know which rule allowed it," you no longer need to ship and parse decision logs to answer the follow-up question — what did that rule claim about itself? The label schema is yours to define; use it.

Upgrade Runbook: Find Your YAML and Your 1.21.0s

This is a minor release, so most fleets will take it via their distribution (Gatekeeper bundles its own OPA — check the Gatekeeper release notes for its cadence) or their base image. For teams running standalone OPA, the binary swap is trivial; the migration is the YAML estate and the pin audit:

# 1. Find YAML that relies on 1.1 booleans (data dirs, config, values):
grep -rInE ':\s*(yes|no|on|off|y|n)\s*$' policies/ config/ data/
#   → every hit is a silent type change; quote it or convert to true/false

# 2. Compile-check every policy against the new type checker (1.21.1):
opa check --strict policies/         # exits 1 on the new rego_type_errors

# 3. Audit for the five-day 1.21.0 window (Sep 24-29, 2026):
#    grep base-image pins across Dockerfiles / CI / helm values
grep -rn --include='Dockerfile*' --include='*.yaml' --include='*.yml' \
    -E '(openpolicyagent/opa|.*opa[^[:alnum:]])[:@v]?1\.21\.0' .
#    anything that resolved 'opa:1.21' to 1.21.0 during the window needs a
#    rebuild on 1.21.1 — do not trust cached layer digests

# 4. Binary upgrade (standalone server), sha256-verified:
curl -fLO https://github.com/open-policy-agent/opa/releases/download/v1.21.1/opa_linux_amd64
curl -fLO https://github.com/open-policy-agent/opa/releases/download/v1.21.1/opa_linux_amd64.sha256
sha256sum -c opa_linux_amd64.sha256      # 23c654cd…f48fbb
install -m 0755 opa_linux_amd64 /usr/local/bin/opa
opa version

One more note for CI: if you lint Rego, the Regal linter is the community-standard companion tool, and its nightly compatibility run against OPA main is now part of OPA's own CI — a good sign for catching the next 1.21.0-class regression before it ships.

Verdict

Upgrade — to 1.21.1, after a YAML audit, not before. The indexer rewrite is the best performance news OPA has shipped in years: a ~660x eval-time reduction on a shape that production policies actually have, with correctness intact and the project's own docs updated to match. The two breaking changes are both defensible engineering — YAML 1.2 is where the ecosystem already was, and empty-literal typing makes the compiler catch what it previously let through — but neither warns you at upgrade time; the YAML one doesn't even fail a compile, it just quietly changes what yes means. Run opa check against every policy and grep every YAML file that touches OPA before you ship this. And treat 1.21.0 as radioactive: the panic was real, we reproduced it, and the fix shipped in five days — which is both a compliment to the maintainers and a reminder that a .0 release from any project is a bet, not a version.

Credit Where Due

The index rewrite is overwhelmingly the work of Stephan Renatus and Stephen Spaink — #9161, #9164, #9235, #9081, #9190 and the compiler-side type-checker work respectively — with Tobias Stockert on the prefix-indexing PR and Anders Eknert across the AST cleanups and the Regal compatibility work. Johan Fylling shipped the debug stack-trace framing, Charlie Egan modernized the website (local Pagefind search, AI-guidelines updates), and the 1.21.1 hotfix came from the same Stephan Renatus, from a report by @tun0. Community reporters who moved this release: @scnewma, @johnc-c (the YAML trap), @disaverio (#7275), @vlsi, @zscott (#7461), @charlesdaniels, @kmadan, @rothenes, @kishorviswanathan (#6854), @nikpivkin, @stobias123, @msahmi. The upstream issues are 3.5 years old in the YAML case — this is what maintenance on a graduated CNCF project actually looks like.