Skip to content

SYSTEM Cited by 1 source

Grafana Tempo

Definition

Grafana Tempo is Grafana Labs' open-source, horizontally-scalable distributed-tracing backend. It stores trace data as columnar blocks (vParquet) in object storage, queries them with TraceQL (a trace-oriented query language), and derives metrics from spans at ingest via an optional metrics-generator (RED metrics + service graphs). It is the "traces" leg of Grafana Labs' observability stack alongside Loki (logs), Mimir (metrics), and Pyroscope (profiles), surfaced through Grafana as the UI.

Architecturally Tempo follows the same object-storage-as-source-of-truth shape as its siblings (see patterns/tiered-storage-to-object-store): stateless read/write components over immutable blocks in object storage, with background compaction.

Ingestion: Kafka read/write-path split

Tempo 3.0 moved microservices-mode ingestion onto Kafka, splitting the read and write paths: the distributor produces records into Kafka, while block-builders, live-stores, and metrics-generators consume them. Tempo 3.1 hardened this Kafka integration (Source: sources/2026-10-01-grafana-tempo-31-release-kafka-traceql-metrics-trace-redaction):

  • Authentication / transport: the 3.0 client could do only SASL PLAIN over an unencrypted connection. 3.1 adds TLS (and optional mTLS via client cert/key) plus SCRAM-SHA-256, SCRAM-SHA-512, OAUTHBEARER, and AWS_MSK_IAM. Credentials can be env-expanded with -config.expand-env=true.
  • Rack-aware fetching (KIP-392): the client_rack option lets consumers fetch from an in-zone replica instead of always hitting the partition leader, avoiding cross-AZ transfer billing. It is a read-side setting (consumers only, not the distributor) and requires brokers to carry rack IDs. This is a direct concepts/egress-cost optimization — the same KIP-392 follower-fetching lever documented on systems/kafka.
  • Configurable producer compression: producer_compression (none/gzip/snappy/lz4/zstd, default snappy) lets Tempo target Kafka-protocol backends that accept only one codec — e.g. Azure Event Hubs (gzip-only) becomes a usable backend. See concepts/compression-codec-tradeoff.

TraceQL and TraceQL metrics

TraceQL selects spans/spansets by resource.* and span.* attributes. TraceQL metrics (GA in 3.0) compute ad-hoc metrics directly from trace data (rate, count_over_time, quantile_over_time, compare, …). Tempo 3.1 extended it:

  • Sampling-aware extrapolation — the experimental with(extrapolate=true) hint reads the per-span sampling rate that OpenTelemetry probability-sampling-spec samplers stamp onto the span's tracestate field and scales each span accordingly (at 50% sampling each stored span counts as 2). Spans with no recorded rate count as 1, so partial SDK adoption is safe. Nothing new is written to blocks; queries without the hint ignore tracestate. Mirrors the metrics-generator's enable_tracestate_span_multiplier so ad-hoc queries agree with pre-aggregated tempo_spanmetrics_* series. Requires vParquet4+. Applies to rate, count_over_time, sum_over_time, avg_over_time, histogram_over_time, quantile_over_time, compare — but not min/max_over_time.
  • Arithmetic between metric queries — + - * / combine two metric queries, so error rate is one query: ({ status = error } | rate() by (resource.service.name)) / ({ } | rate() by (resource.service.name)). Label sets must match exactly to combine; a label-less side broadcasts. A single with(...) hint at the end applies to both sub-queries.
  • Span-only fetch read path — queries that don't need full trace structure process individual spans instead of whole traces. Experimental in 3.0; default for vParquet5 blocks (what 3.1 writes). Roughly 2× faster on simple metrics queries in a warmed-up test; opt-out per tenant/query. Rooted in columnar block layout — see concepts/columnar-storage-format.

Metrics-generator (RED metrics + service graphs)

The optional metrics-generator derives metrics from ingested spans: RED (Rate/Error/Duration) metrics and service graphs (edges built by pairing client and server spans for a connection). 3.1 improvements:

  • Service-graph diagnostics: an unmatched_span_kind label on the expired-edges counter reveals which side of a missing edge never arrived (e.g. many SPAN_KIND_SERVER expirations ⇒ server-side instrumentation that never resolves into a node).
  • OTel convention tracking: recognizes both db.system and the renamed db.system.name (OTel semantic conventions v1.30.0), so an SDK upgrade no longer silently drops database nodes from the service map (db.system takes precedence when both are present).
  • Zero-allocation hot path: span-metrics and service-graphs processors borrow pooled label buffers from the registry instead of allocating per span. Steady-state service-graphs benchmarks report 0 allocations per operation, with no change to metric names/labels/values — just lower generator CPU/memory. The generator touches every span, so per-span label-set work (the dominant cost — see concepts/metric-cardinality) sets the floor.

Trace redaction over immutable blocks

Tempo 3.0 added trace redaction to permanently remove leaked PII (emails, auth tokens, account numbers) from span attributes without waiting for retention. 3.1 made it query-driven (Source: sources/2026-10-01-grafana-tempo-31-release-kafka-traceql-metrics-trace-redaction):

  • tempo-cli redact --query '{span.attribute = "<leaked PII>"}' replaces the old enumerate-every-trace-ID workflow. --dry-run first — Tempo counts what it would remove without touching blocks and prints a batch ID + job count; re-run without --dry-run to rewrite blocks (irreversible).
  • Accepted syntax is deliberately tiny — a single spanset filter with equality comparisons on resource.*/span.*, combined with &&/|| — because a wrong query permanently deletes data.
  • Redaction rewrites immutable object-storage blocks (see concepts/immutable-object-storage) and pauses per-tenant compaction while running. --start/--end (accepting now, now-7d, RFC3339) scope the job; compaction catches up between runs. Version-skew hazard: the time window is only honored once every scheduler and worker in the cell runs ≥3.1 — an older worker ignores the window and removes every match in each block, with no error and no recovery.

Read-time span pruning

The trace-by-ID v2 endpoint can collapse groups of similar leaf spans into one summary span at read time (?span_pruning=true, or default via span_pruning_enabled_by_default; also settable per tenant). The summary span carries aggregation.span_count plus min/max/avg/total duration — a request with 1 root + 100 near-identical DB spans (101 total) returns 2 spans in a smaller response. Spans group only when name, kind, status, and parent name match, so an error is never folded in with successes. Happens at read time; nothing leaves storage, and an explicit span_pruning on a request overrides the default (so a caller needing the full trace can still ask). Trace-by-ID v2 also gained TraceQL filter support (?q={ span.http.status_code = 500 }), with keep_hierarchy=true to also return each match's ancestor path to the root. This also covers the Grafana UI and the Tempo MCP server's get-trace tool (neither sends the param yet). Experimental — params/format may change.

Block versions

  • vParquet4 — minimum for with(extrapolate=true) sampling extrapolation.
  • vParquet5 — written by Tempo 3.1; enables span-only fetch by default.
Last updated · 766 distilled / 2,225 read