Skip to content

Tempo 3.1: Kafka ingestion hardening, sampling-aware TraceQL metrics, query-driven trace redaction, span pruning

Summary

Grafana Tempo 3.1 is a feature release for the open-source distributed-tracing backend, but under the release-note framing it carries real distributed-systems engineering across four areas. (1) Kafka ingestion — Tempo 3.0 moved microservices-mode ingestion onto Kafka; 3.1 adds TLS/mTLS and four SASL mechanisms (SCRAM-SHA-256/512, OAUTHBEARER, AWS_MSK_IAM), rack-aware fetching (KIP-392) to cut cross-AZ transfer cost, and configurable producer compression codec so Kafka-protocol backends like Azure Event Hubs (gzip-only) become usable. (2) TraceQL metrics gains sampling-aware extrapolation — an experimental with(extrapolate=true) hint reads the per-span tracestate sampling rate and scales counts back to true traffic — plus arithmetic between metric queries (error rate = errors / total in one query) and a span-only fetch read path (now default on vParquet5) that runs simple metrics queries ~2× faster. (3) Trace redaction can now target spans by a TraceQL query (not just enumerated trace IDs), running as per-tenant batch jobs that rewrite immutable object-storage blocks and hold compaction off during the run. (4) Span pruning collapses groups of similar leaf spans at read time into one summary span, and the trace-by-ID v2 endpoint gains TraceQL filter support. The post is a community-contribution-heavy release but the Kafka egress/compression and sampling-extrapolation sections are genuine systems design.

Key takeaways

  • Kafka client gained TLS + 4 SASL mechanisms. Tempo 3.0 shipped a Kafka client that could authenticate only with SASL PLAIN over an unencrypted connection — a TLS-requiring broker would refuse it. 3.1 adds TLS (and optional mTLS via client cert/key) plus SCRAM-SHA-256, SCRAM-SHA-512, OAUTHBEARER, and AWS_MSK_IAM. Config supports -config.expand-env=true for env-var interpolation of credentials. (Source: this article)
  • Rack-aware fetching (KIP-392) is an egress-cost optimization. Tempo always fetched from the partition leader wherever it lived; when the leader sits in another AZ, every trace read crosses a zone boundary and the cloud provider bills the transfer. The new client_rack option lets consumers fetch from an in-zone replica (follower fetching). It is a read-side setting — applies to block-builders, live-stores, and metrics-generators, not the distributor writing records — and requires brokers to have rack IDs configured. (Source: this article) → see concepts/egress-cost, systems/kafka.
  • Producer compression codec is now configurable. The distributor compresses every batch written to Kafka; until 3.1 the codec was fixed (defaulting to snappy). The new producer_compression option (none/gzip/snappy/ lz4/zstd) unblocks backends that accept only one codec — e.g. Azure Event Hubs speaks the Kafka protocol but accepts only gzip. (Source: this article) → see concepts/compression-codec-tradeoff.
  • Sampling-aware metrics via tracestate. If you tail-sample 50% of traces, a naive rate() reports half the real traffic because Tempo counts only surviving spans. The experimental { } | rate() with(extrapolate=true) hint reads the sampling rate that OpenTelemetry probability-sampling-spec samplers stamp onto each span's tracestate field, and scales each span accordingly (at 50% sampling each stored span counts as 2, so 1,000 matched spans report 2,000). Spans with no recorded rate count as 1 — so partial adoption is safe. Nothing new is written to blocks (tracestate is already stored); queries without the hint don't read it. Requires vParquet4 blocks or later. Mirrors the metrics-generator's existing enable_tracestate_span_multiplier so ad-hoc queries agree with pre-aggregated tempo_spanmetrics_* series. Applies to rate/count_over_time/sum_over_time/avg_over_time/ histogram_over_time/quantile_over_time/compare; not to min/max_over_time (sampling can't change the extreme value actually observed). (Source: this article) → see systems/opentelemetry.
  • Arithmetic between metric queries. + - * / now work between two TraceQL metric queries, so error rate is a single query: ({ status = error } | rate() by (resource.service.name)) / ({ } | rate() by (resource.service.name)). Series combine only when label sets match exactly; a label-less side is broadcast. One with(extrapolate=true) at the end of the expression applies to both sub-queries (putting it inside a sub-query is a syntax error). (Source: this article)
  • Span-only fetch read path, default on vParquet5. Metrics queries that don't need full trace structure now process individual spans instead of whole traces, cutting latency and memory. Introduced experimental in 3.0; now default for vParquet5 blocks (what Tempo 3.1 writes). In one warmed-up test { } | rate() completed in slightly over half the time with span-only fetch enabled. Opt-out per tenant or per query. (Source: this article) → see concepts/columnar-storage-format.
  • Query-driven trace redaction. tempo-cli redact previously accepted only trace IDs, forcing you to enumerate every affected trace. 3.1 accepts a TraceQL query (--query '{span.attribute = "<leaked PII>"}'). Run --dry-run first: Tempo counts what it would remove without touching blocks and prints a batch ID + job count; re-run without --dry-run to rewrite blocks (irreversible). Accepted syntax is deliberately tiny — a single spanset filter with equality comparisons on resource.*/span.*, combined with &&/|| — because a wrong query permanently deletes data. (Source: this article) → see concepts/immutable-object-storage.
  • Redaction holds compaction off per-tenant. Redaction rewrites immutable object-storage blocks, and Tempo pauses compaction for a tenant while a redaction runs; --start/--end (accepting now, now-7d, or RFC3339) scope the job to run faster in high-volume installs, with compaction catching up between runs. Version-skew hazard: --start/--end are only safe once every scheduler and worker in the cell runs ≥3.1 — an older worker ignores the window and removes every match in each block regardless of timestamp, with no error and no recovery. (Source: this article)
  • Span pruning collapses repetitive leaf spans at read time. A request that fires 100 near-identical DB queries yields 1 root + 100 look-alike spans. The trace-by-ID v2 endpoint (?span_pruning=true, or default via span_pruning_enabled_by_default) collapses each group of similar leaf spans into one summary span carrying aggregation.span_count plus min/max/avg/total duration — 101 spans become 2, in a smaller response. Spans group only when name, kind, status, and parent name match, so an error is never folded in with successes. Happens at read time; nothing leaves storage. Trace-by-ID v2 also gained TraceQL filter support (?q={ span.http.status_code = 500 }), with keep_hierarchy=true to return each match's ancestor path. (Source: this article)
  • Service-graph diagnostics + OTel convention tracking. An unmatched_span_kind label on the expired-edges counter now tells you which side of a missing service-graph edge never arrived (e.g. many SPAN_KIND_SERVER expirations ⇒ server-side instrumentation that will never resolve into a node). Tempo 3.1 also recognizes both db.system and the renamed db.system.name (OTel semantic conventions v1.30.0) for identifying DB calls, so an SDK upgrade no longer silently drops database nodes from the service map. (Source: this article)
  • Zero-allocation metrics-generator hot path. The metrics-generator processes every ingested span, so per-span work sets the cost floor; building label sets was a big share. The span-metrics and service-graphs processors now borrow pooled label buffers from the registry instead of allocating per span, and steady-state service-graphs benchmarks report zero allocations per operation — same metric names/labels/values, lower generator CPU and memory. (Source: this article) → classic hot-path allocation elimination; see concepts/metric-cardinality for why label work dominates generator cost.

Systems / concepts extracted

Operational numbers

  • Four new SASL mechanisms added (SCRAM-SHA-256, SCRAM-SHA-512, OAUTHBEARER, AWS_MSK_IAM) on top of the pre-existing PLAIN.
  • Default producer compression codec: snappy.
  • Span-only fetch: { } | rate() ran in slightly over half the time vs disabled, in a warmed-up test (~2× speedup on simple metrics queries).
  • Extrapolation example: at 50% sampling, 1,000 matched spans report 2,000.
  • Span pruning example: 1 root + 100 similar DB spans (101 total) → 2 spans.
  • Service-graphs processor steady-state benchmark: 0 allocations per operation after pooled-buffer change.
  • Experimental features gated on block version: with(extrapolate=true) requires vParquet4+; span-only-fetch default on vParquet5.

Caveats

  • Several features are explicitly experimental and may change: the with(extrapolate=true) hint, span pruning (span_pruning params/format), and the span-only fetch opt-outs.
  • Redaction is irreversible and the --start/--end time window is only safe once all schedulers/workers in a cell run ≥3.1 — a mixed-version fleet silently over-deletes. Redaction also pauses per-tenant compaction while running.
  • Sampling extrapolation depends on OpenTelemetry's probability-sampling spec, which is still in development and not supported in every language SDK yet; spans without a recorded rate count as 1 (safe but under-counts those paths).
  • This is a release/announcement post (Tier-2 Grafana). It is heavy on community-contribution shout-outs and config snippets; included because the Kafka egress/compression, sampling-extrapolation, and read-time pruning sections carry genuine architectural reasoning well above the ~20% threshold.

Source

Last updated · 766 distilled / 2,225 read