Skip to content

SYSTEM Cited by 5 sources

OpenTelemetry

Definition

OpenTelemetry (OTel) is a vendor-neutral, CNCF-hosted standard and set of SDKs/collectors for generating, collecting, and exporting telemetry — traces, metrics, and logs — from applications and infrastructure. Its value is portability: instrument once against the OTel APIs and export to any compatible backend (Amazon CloudWatch, Prometheus/Grafana, Datadog, …) without rewriting instrumentation per vendor.

Role in stateless MCP deployments

In the AWS reading of MCP 2026-07-28, OpenTelemetry is the standardized backbone for the observability the protocol now assumes (Source: sources/2026-09-01-aws-mcp-went-stateless-is-your-aws-mcp-server-deployment-well-architected):

  • Tracing. Requests carry W3C Trace Context in _meta, so a tool call traces end to end through any OpenTelemetry-compatible backend, including CloudWatch.
  • Logging. MCP's proprietary protocol logging is deprecated in favor of stderr and OpenTelemetry — structured logging through standardized, queryable formats rather than a bespoke protocol channel.

The through-line: the protocol stops inventing its own observability channels and defers to the OTel/W3C standards, so ordinary observability infrastructure can operate an MCP fleet without becoming MCP-aware.

Seen in

  • sources/2026-09-01-databricks-how-we-eliminated-1-million-a-year-of-wasted-ai-agent-spend — Unity Gateway emits one OTel trace per MCP tool invocation (tool name, args, error, tokens, latency, session ID) into a single unified trace table. Querying it with Genie One surfaced seven silent tool bugs (~$1.2M/year) — OTel as the substrate for trace-driven tool-failure diagnosis.
  • sources/2026-09-30-cloudflare-detect-and-send-production-issues-straight-to-your-agent — OTel as the app-context enrichment API baked into a serverless runtime. Cloudflare Workers exposes a built-in OpenTelemetry API (import { tracing } from "cloudflare:workers"; tracing. getActiveSpan()?.setAttribute("user.id", …)) with no package to install. Because platform-captured errors know what failed but not who it hit, teams attach user.id / account.id / session.id as span attributes; those identifiers then appear on every occurrence grouped into a Workers Issue, so failures can be localized to one account or session before the issue is handed to a coding agent. OTel as the "add the business identifiers the platform can't know" seam in error monitoring.
  • sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees — OTel traces as the cost-comparison substrate for model rollout. Unity Gateway logs all traces centrally with cost, letting Databricks compare a pilot cohort's per-session spend on the prior model generation vs the new one. Because early adopters try harder problems with new models, the raw $/session average is misleading, so sessions are stratified (single- vs multi-turn × file-edits) and re-weighted before comparison — OTel cost data as one of three promote/drop signals (patterns/experimental-tier-model-promotion).
  • sources/2026-10-01-grafana-tempo-31-release-kafka-traceql-metrics-trace-redaction — OTel probability sampling + semantic conventions as the correctness substrate for trace-derived metrics. Grafana Tempo 3.1's experimental with(extrapolate=true) TraceQL hint reads the sampling rate that OTel probability-sampling-spec samplers stamp onto each span's tracestate field and scales the span to reconstruct true traffic (at 50% sampling, each stored span counts as 2). Spans with no recorded rate count as 1, so partial SDK adoption is safe — a direct demonstration of why standardizing the sampling signal in the span matters. Tempo 3.1 also tracks OTel semantic-conventions v1.30.0, accepting both db.system and the renamed db.system.name so an SDK upgrade doesn't silently drop database nodes from service-graph maps.
Last updated · 766 distilled / 2,225 read