OpenTelemetry Collector v0.162: The queuebatch Rename, a Log-Dropping Bug Nobody Benchmarked, and Keepalive Defaults That Flip Overnight
Sources
- OpenTelemetry Collector v0.162.0 release notes
- PR #16018 — rename queuebatch processor to queue_batch
- Issue #14396 — Rename components to match naming convention
- PR #15937 — Drop only the oversized item when splitting a batch
- Issue #15936 — split() drops subsequent valid records on oversized item
- PR #16026 — pkg.confighttp.PrioritizeNewKeepalive feature gate
- Issue #9380 — Stabilize module confighttp (keepalive refactor parent)
- Issue #15894 — partitionIdleCycles too aggressive, scrape churn
- PR #15683 — preserve status code for unsupported Content-Type
- Issue #15995 — errorHandler returns 500 on malformed Content-Type
- PR #15837 — otelcol_memorylimiter_refused_requests metric
- PR #15994 — mdatagen context-propagation lifecycle tests
- Issue #15747 — unbounded memory growth in queuebatch split path (still open)
- Issue #15544 — pprofile dictionary merge O(N·L²) fix
- Issue #14526 — bounded multi-batching partition LRU
- OpenTelemetry Collector binary releases (v0.162.0 distributions)
- OpenTelemetry Collector documentation
- Collector Builder (ocb) documentation
OpenTelemetry Collector core v1.68.0/v0.162.0 shipped September 28, 2026, and its changelog reads like a Tuesday: a processor rename, a config-key migration, some performance work. Three of those items deserve more attention than the release title gives them, because one of them is a data-loss bug in the batching path that logs pipelines could hit for months without noticing, one is a default flip in HTTP connection behavior across every receiver and exporter in your fleet, and one is a rename that will fail your Collector at startup — no deprecation warning, no alias, no grace period. This is the analysis for operators who ship otelcol to production, not people who collect release notes.
Context for versioning newcomers: the open-telemetry/opentelemetry-collector repo (the core, where this release lives) publishes as two coordinated numbers — the semantic core API version (v1.68.0) and the component-carrying distribution version (v0.162.0). The prebuilt otelcol and otelcol-contrib binaries come from the opentelemetry-collector-releases repo, whose v0.162.0 binaries followed a day later. If you build your own distribution with the Collector Builder (ocb) — and most serious deployments do — the rename in this release is the first thing your build manifest will reject.
The Rename That Fails Your Build: queuebatch → queue_batch
PR #16018, authored by @andrzej-stencel and merged September 24, renames the queuebatch processor component to queue_batch, closing issue #14396 ("Rename components to match naming convention"). The Go package name stays processor/queuebatchprocessor; what changes is the component type: you write in YAML and the name ocb resolves in your builder manifest. There is no compatibility alias, and that is deliberate: the PR body points at the project's own component-naming guideline and notes the component is still Development stability, so the maintainers chose not to preserve the old name.
Before anyone shrugs: the queue_batch processor is not a toy. It re-exports the exact same Config struct as the exporter-side queue helper (exporterhelper.QueueBatchConfig — verified in processor/queuebatchprocessor/config.go in the v0.162.0 tree), which means it gives you the full partitioned-batching engine — metadata_keys partitioning, LRU partition cache, byte-sizing — as a processor, in front of any exporter, including ones with no native partition support. If you were using it to partition logs by tenant before a fan-out exporter, your pipeline definition now contains a component name that v0.162.0 refuses to start. The migration is one line:
# v0.161.0 and earlier # v0.162.0
processors: processors:
queuebatch: queue_batch:
batch: batch:
flush_timeout: 200ms flush_timeout: 200ms
min_size: 8192 min_size: 8192
service: service:
pipelines: pipelines:
logs: logs:
processors: [queuebatch] processors: [queue_batch]The second breaking change rides the same rename. mdatagen — the tool that generates the boilerplate every component ships — now generates lifecycle tests that assert the component propagates the incoming context unchanged to the next consumer by default (PR #15994, by @TylerHelmuth, merged September 24). That is a correctness bar, not an API break for users — but it changes what "my custom processor passes the test suite" means, and it has already claimed its first documented victim: the queue_batch processor itself. Its metadata.yaml in the release tree sets skip_context_propagation: true with the comment that it "buffers data and flushes it on a detached context, so it cannot propagate the incoming context to the next consumer." If you maintain internal or contrib processors, expect your CI to fail on this assertion until you either fix the propagation or consciously opt out with the same toggle.
Practical blast radius of the rename: the prebuilt otelcol-contrib binary does not include queue_batch at all (its v0.162.0 manifest lists 33 processors; none is queue_batch — the component is Development-stability and ships only in custom ocb builds). So the broken configs are exclusively in hand-built distributions and Helm-rendered values that name the processor explicitly. Grep your fleet for the literal string queuebatch before you roll the new version — a failed component name is a Collector that exits at config parse, which in a DaemonSet is a CrashLoopBackOff you could have caught with one grep.
The Bug That Dropped Your Logs: Oversized Item, Entire Batch Gone
The most operationally important change in this release is buried under "Bug fixes" with no drama: PR #15937 (authored by @avadla, merged September 28 — the day the release branch cut), fixing issue #15936: "exporterhelper/queuebatch: split() drops subsequent valid records when an oversized item is encountered."
The mechanism is worth understanding, because it explains why nobody benchmarked this bug out of existence — it only fires under load, and it fails silent. When a queued batch exceeds the exporter's limits (an OTLP payload larger than max_size, a tenant whose log line swallowed a stack trace, a metrics batch over the byte budget), the queue helper splits it into chunks. Pre-0.162.0, the split logic hit an oversized item and dropped every item queued behind it in the same split pass — not the oversized item alone, the whole tail of the batch. The release notes state the fix plainly: "Drop only the oversized item when splitting a batch, instead of discarding every item queued behind it." Note the scope-limiting sentence that follows: "Applied to logs only for now."
queue: [rec_1 .. rec_49, OVERSIZED_rec_50, rec_51 .. rec_200]
│
exporter limit hit: batch must be split
│
┌──────────────────────┴─────────────────────────────┐
│ pre-v0.162.0 split() │
│ │
│ chunk A: rec_1 .. rec_49 SHIPPED ✓ │
│ chunk B: OVERSIZED_rec_50 DROPPED ✗ (correct) │
│ chunk C: rec_51 .. rec_200 DROPPED ✗ (the bug) │
│ │
│ v0.162.0 (#15937): │
│ chunk A: rec_1 .. rec_49 SHIPPED ✓ │
│ OVERSIZED_rec_50 DROPPED ✗ (correct) │
│ chunk C: rec_51 .. rec_200 SHIPPED ✓ (fixed) │
└──────────────────────────────────────────────────────┘
failure signature: logs pipelines silently lose up to a
batch-worth of records behind any oversized item —
no error, no drop metric, gaps in every affected tenantWhy is this release's quiet headline? Because the failure mode is invisible from the sending side. The receiver accepted the data (it is in the queue), the exporter reports no error (it shipped something), and the only evidence is a missing log among the ones that did ship. If you run log pipelines through partitioned or size-limited exporter queues — the default shape for multi-tenant gateways — and you have ever traced "a log line that should be there isn't," this bug is a candidate explanation, and v0.162.0 is your fix. The honest caveat is scope: the fix lands on the logs path first; the same split machinery serves traces and metrics, so watch the 0.163 changelog before assuming parity. And the deeper memory-growth issue in the same split path — issue #15747, where a slow exporter starves flush workers and the split path accumulates unbounded live chunks — remains open at release time, with its candidate fix (PR #15798) still unmerged. The scale is documented in that PR's problem description: heap snapshots above 36GB, production RSS exceeding 200GB, with every live chunk uncounted by queue_size or any other bound. Size your collector memory limits assuming #15747 is still true.
Keepalive Defaults Flip — and Your Old Config Keys Stop Meaning What They Said
The change with the widest fleet impact is the one nobody's changelog summary gets right. PR #16026 (authored by @codeboten, merged September 25) enables a new feature gate, pkg.confighttp.PrioritizeNewKeepalive, at beta stage, on by default, registered as of v0.162.0 (verified in the generated feature-gate registry in the release tree). The gate's contract, from its own registration description: "When enabled, the Keepalive configuration is prioritized over the deprecated configuration fields." It exists to unwind a config mess that dates back to issue #9380 — "Stabilize module confighttp" — open since January 2024, and it changes what your existing HTTP blocks mean.
Here is the semantics trap in plain terms. Every HTTP-based receiver and exporter in the Collector accepts four flat keepalive-adjacent fields, deprecated since v0.160.0: idle_conn_timeout, max_idle_conns, max_idle_conns_per_host, disable_keep_alives (client side), plus idle_timeout (server side). Until now, those flat fields were the source of truth. In v0.162.0 with the gate on, a new nested keepalive: block takes precedence when you set it — and the docs are explicit that "keep-alives are always enabled" unless you explicitly disable them via the new keepalive::enabled: false (server schema, verified in the generated config schema). The defaults inside the new block mirror Go's net/http transport defaults — idle_conn_timeout: 90s, max_idle_conns: 100 — which for most fleets is what the old flat fields already produced. The risk case is the config that sets both the old flat fields and a new keepalive block, or a config managed by a Helm chart that renders both shapes across versions: the nested block now wins, silently. The escape hatch is turning the gate off in service.telemetry.feature-gates, restoring the old precedence until you finish the migration.
# receivers & exporters with an HTTP server side:
receivers:
otlp:
protocols:
http:
endpoint: 0.0.0.0:4318
# NEW in v0.162.0, takes precedence over deprecated flat fields:
keepalive:
enabled: true # default: true; 'false' is the ONLY way to disable now
idle_timeout: 60s # server side; 0 falls back to read_timeout
exporters:
otlphttp:
endpoint: https://otlp.example.com
keepalive:
idle_conn_timeout: 90s # default: 90s (mirrors net/http)
max_idle_conns: 100 # default: 100; 0 = no limit
max_idle_conns_per_host: 0 # 0 -> net/http.DefaultMaxIdleConnsPerHost
# Rollback pin if the precedence change breaks a chart-rendered config:
service:
telemetry:
feature-gates:
- pkg.confighttp.PrioritizeNewKeepalive # '-' prefix disables itWhy should a platform engineer care about keepalive plumbing? Because connection churn on the OTLP ingestion path is a real, measurable cost — TIME_WAIT sockets on high-cardinality gateway fleets, TLS handshakes multiplied across every exporter scrape, load balancers whose idle timeouts must be tuned against idle_timeout to avoid mid-request kills. The keepalive unification is the config-level prerequisite for the confighttp stabilization track that issue #9380 tracks, and the gate's stated intent ("removed in a few releases", per the PR body) means the old flat fields are now on a countdown. Migrate your charts this quarter; do not be the pager recipient when the removal release lands.
Partitioned Batching Grows Knobs — and Finally Respects Your Scrape Interval
If you use sending_queue::batch::partition on exporters (or the queue_batch processor, which shares the config), this release fixes an annoyance you have probably already worked around. Pre-0.162.0, the idle eviction of a partition batcher was hardcoded: ten flush cycles with no data, and the partition is torn down — a literal partitionIdleCycles = 10 constant with a TODO make this configurable comment sitting next to it in the v0.161.0 source. Issue #15894 titled it precisely: "partitionIdleCycles is too aggressive so every metric scrape churns its shard." With a default 200ms flush timeout, ten cycles is two seconds — a Prometheus scrape arriving every 30 or 60 seconds re-created its partition batcher on virtually every scrape.
v0.162.0 replaces cycle-counting with wall-clock time and makes it configurable: idle_timeout under batch::partition, default 90 seconds — the source comment is explicit that this is "large enough to keep a partition alive across common metrics scrape intervals (up to 60s)." Alongside it, issue #14526 lands the other half: the multi-batcher partition LRU cache is now bounded and observable — cache_size (default 10000, must be positive) caps active partition batchers, evicting least-recently-used partitions by flushing them, and exports two new gauges: otelcol_exporter_queue_batch_partition_cache_size and otelcol_exporter_queue_batch_partition_cache_capacity. For multi-tenant gateways where unbounded per-tenant partition accumulation was a silent memory leak class, that bound plus its gauges is the difference between an alert and a 3 a.m. heap dump.
exporters:
otlphttp:
endpoint: https://otlp.example.com
sending_queue:
num_consumers: 10
queue_size: 10000
batch:
flush_timeout: 200ms
min_size: 8192
partition:
metadata_keys:
- tenant.id # one batcher per distinct value combination
idle_timeout: 5m # NEW: was hardcoded 10 flush cycles (~2s)
# keep ABOVE your arrival interval
cache_size: 2000 # NEW: LRU cap on live partition batchers
# default 10000; evicted partitions are flushedSmaller Changes That Punch Above Their Changelog Line
| Change | What actually happens | Why you care |
|---|---|---|
| New memory-limiter refusal metric (PR #15837) | The memory-limiter extension now exports otelcol_memorylimiter_refused_requests, counting requests it rejected, by transport (HTTP and gRPC). The processor variant had refusal counters; the extension mode had none — refusals were invisible. | If you run the extension (the recommended pattern for the receiver tier), you can finally alert on "the collector is refusing telemetry" instead of inferring it from downstream gaps. The metric is marked [Alpha]. |
| OTLP HTTP error codes stop lying (PR #15683, fixes #15995) | Previously, a request with a malformed or unsupported Content-Type — including a 401 from a server-side auth extension — fell through to 500 {"code":13,"message":"failed to marshal error message"}. Now the real status code is preserved: a client error stays a 4xx, only server faults are 5xx. Extends the earlier missing-Content-Type fix (#13414). | Your SLO dashboards stop counting misconfigured senders as collector outages, and your on-call stops paging for other teams' auth failures. Retry logic in senders (which keys on 4xx vs 5xx) also stops hammering a healthy collector. |
| Profiles dictionary merge: O(N·L²) → O(N+L) (issue #15544) | The profile-data dictionary merge during exporter-queue batching replaced a linear-scan dedup with an indexed lookup. Release notes: merged contents and remapped indices identical — pure CPU win, no behavior change. | If you are ahead of the curve on continuous profiling ingestion (OTel profiles signal), large profile payloads stopped quadratic-ing your exporter CPU. Everyone else: free headroom. |
| Receiver allocation diet (issue #15998) | Long-lived-context receivers no longer build a span Link that the SDK always discarded when no parent span context exists. Removes 3 of 7 allocations per receive operation. | The hot path of every OTLP receiver gets measurably cheaper. Multiply by your requests-per-second and enjoy the GC pause you are not having. |
| xhash map hashing de-quadratic-ified (issue #15990) | xhash.MapHash no longer takes quadratic time in the number of map entries. | Attribute-heavy telemetry with many map entries stops punishing you at hash time. Small code, wide blast radius across all signals. |
The Upgrade Playbook
# 1. Find configs that name the OLD processor (startup failure if missed)
grep -rn "queuebatch" charts/ manifests/ rendered-configs/
# → every hit must become queue_batch BEFORE the new binary rolls
# 2. Find keepalive flat fields (deprecated, now shadowed by keepalive:)
grep -rnE "idle_conn_timeout|max_idle_conns|disable_keep_alives" charts/
# → migrate into keepalive: blocks, or pin the feature gate OFF this cycle
# 3. Decide partition idle_timeout per pipeline from arrival intervals
# interval 30s → idle_timeout >= 90s (default) is fine
# interval 60s → default 90s is fine (that is WHY it is 90s)
# bursty/hourly → set idle_timeout ABOVE the longest gap, or accept churn
# 4. Cap partition cache_size on multi-tenant gateways and wire the gauges
# otelcol_exporter_queue_batch_partition_cache_size vs _capacity
# 5. Alert on the new refusal metric (extension mode)
# otelcol_memorylimiter_refused_requests > 0 means the collector
# is dropping work to save its own heap — find the memory pressureHidden Costs and Breaking Changes
- The rename has no grace period. Development-stability components get no alias; a config naming
queuebatchfails at parse time on v0.162.0. Since queue_batch ships only in custom ocb builds (verified: not in the otelcol-contrib v0.162.0 manifest), every fleet running it is a fleet with a build manifest to fix too. - The log-drop fix is logs-only. #15937's fix is applied to the logs split path; the changelog says so explicitly. Traces and metrics keep the old drop-the-tail behavior until a follow-up — plan your "we lost batch tails" forensics accordingly for non-log pipelines.
- The split-path memory-growth bug (#15747) is still open. 36GB heaps, 200GB RSS in production reports, candidate fix unmerged at release time. The 90s idle timeout and cache_size bound help, but byte-sized queues with slow exporters are still exposed. Memory limits on collector containers are not optional this cycle.
- Keepalive precedence is a chart-compatibility hazard. Any tooling that renders both the deprecated flat fields and a new keepalive block will silently change behavior in favor of the nested block. Pin the gate off until your values are consistent, then migrate.
- Go 1.26 is the build floor. The processor module's go.mod in the release tree declares go 1.26.0 — if your custom distribution pins an older toolchain in CI, expect the upgrade PR to be bigger than the component bumps.
- mdatagen's new context-propagation tests will fail your custom components' CI until each component either propagates the incoming context or sets
skip_context_propagation: truewith a reason. That is a good forcing function, and an afternoon of work you did not budget.
Verdict
Upgrade this cycle if you run log pipelines through size-limited or partitioned exporter queues — the split-path fix (#15937) is a data-integrity repair for a bug that fails silent, and the partition idle_timeout + cache_size knobs close out a config gap (hardcoded 10-cycle eviction) that the source itself carried as a TODO. For the deeper architectural picture of where collectors fit in a telemetry pipeline — agent vs gateway tiers, fan-out topologies — read our OpenTelemetry Collector architecture guide. Everyone else: this is a routine minor whose breaking edge is narrow and greppable — do the two greps (queuebatch, keepalive flat fields), roll on your next maintenance window, and take the free CPU wins. Hold the upgrade back only if you are mid-incident on collector memory and cannot tolerate behavior shifts; note the biggest known memory bug is NOT fixed here, so that is not a reason to stay — size your limits and move. Building the operational posture around this — refusal-metric alerting, partition-capacity dashboards, config-key audits on every minor — is genuinely a platform-engineering function; if you would rather outsource it, the managed tier exists at Grafana Cloud and Amazon Managed Service for Prometheus (with vendors like Datadog consuming OTLP directly), and the self-managed path starts at the Collector docs and the CNCF landscape for every component in between.
Credit Where Due
Release assembled under the project's chloggen discipline, but written by people: @andrzej-stencel for the queue_batch rename and the naming-convention cleanup behind it (#14396, #16018), @avadla for the log-split data-loss fix (#15937) — the release's most important change, landed the day the branch cut — @codeboten for the keepalive precedence gate (#16026) on the confighttp stabilization track, @TylerHelmuth for the mdatagen context-propagation test generation (#15994), @chethanm99 for the memory-limiter refusal metric by transport (#15837, closing @tank-500m's report #15561), and the performance trio of @lazureykis (pprofile dictionary merge, #15544), @paulojmdias (receiver link-allocation diet, #15998), and @IshwarKanse (xhash de-quadratic-ing, #15990). Issue authors did the invisible work of precise bug reports: @dmitryax named the idle-cycle churn that became idle_timeout (#15894), @vengalraoguttha asked for bounded multi-batching (#14526), @boqu caught the error-handler status-code mangling (#15995), and the maintainers carry the confighttp stabilization debt (#9380) that this release moves forward. Code owners for the renamed processor: @jmacd and @iblancasa.