Skip to content

SYSTEM Cited by 1 source

Databricks AI SRE

Definition

AI SRE is Databricks' AI-powered debugging agent for incident investigation. It begins investigating the moment an incident fires, correlating signals across a fleet of 100s of microservices on 1500+ Kubernetes clusters spanning 70+ regions and three clouds, and guides the on-call engineer through root-cause analysis. Its design thesis, drawn from dozens of on-call interviews, is that context assembly — not root-cause insight — is the bottleneck: gathering the right metric, time window, deploy, and upstream dependency consumes 60–80% of investigation time. (Source: sources/2026-08-24-databricks-how-databricks-uses-ai-to-accelerate-incident-investigation)

Two experiences

  • Automatic triage — kicks off the instant an incident fires, before the engineer opens their laptop. Launches three investigation tracks in parallel and correlates them into one diagnostic summary:
  • Platform health checks — is the cloud infrastructure healthy? are there network issues in the region? are upstream dependencies (databases, queues, shared services, Auth) degraded? This eliminates a large class of red herrings up front.
  • Service-level analysis — pulls logs, metrics, traces for the affected service and its immediate dependencies; examines recent deploys and config changes; flags anomalies relative to baseline ("CPU spiked 3x at 2:47 AM, coinciding with a deployment that changed the batch size").
  • Runbook execution — AI SRE assumes a team-specific persona and runs the team's agentic runbook.
  • Interactive debugging — a UI where engineers ask natural-language follow-ups ("was there anything unusual about Kafka consumer lag in the 10 minutes before this alert?"), request additional signals, and drill into specific time windows or components.

Layered architecture

AI SRE is built as a layered platform so first-party bots and team-owned runbooks share one substrate:

Layer Responsibility
Application Layer Where debugging happens — the platform-level incident triage bot; also where third-party AI tools plug in.
Core Engine Where the intelligence lives — a bot framework for building debugging workflows + orchestration for parallel execution, result correlation, and LLM-powered synthesis. The same platform first-party bots and any team's bots run on.
API Layer Purpose-built Observability API, Deployment API, Alerts API over the primitives — handling auth, rate limiting, data normalization, and guardrails. Turns "raw infrastructure" into "debuggable infrastructure" and decouples the tools above from swapping out an underlying system.
Primitives The raw operational data every investigation depends on: metrics, alerts, logs, release information, code. The systems of record; the layer does not replace them.

Reliability engineering principles

  1. Structured checks before open-ended reasoning — deterministic platform health checks and runbook steps run first; the LLM synthesizes and explains but does not decide what data to gather. See structured-checks-before-llm-reasoning.
  2. Transparency over black-box answers — every conclusion links to the underlying evidence (metric, log line, deploy diff) so engineers can audit, not just trust. See traceable-evidence-for-agent-trust.
  3. Graceful degradation — when confidence is low, AI SRE says so and presents the evidence it gathered, organized by relevance. See concepts/graceful-degradation.

Guardrails for agent access

Giving agents access to observability data required redesigning the API layer, not just opening it. Agents hit endpoints in bursts, run checks in parallel, and never back off on their own, so guardrails were needed to keep agents from overwhelming infrastructure that also powers business-critical alerting.

Operational scale

  • Supports 150+ teams, 250+ weekly active users, 2,000+ investigations per day, saving "several hours of debugging time."
  • Fleet under investigation: 100s of microservices, 1500+ K8s clusters, 70+ regions, 3 clouds.

Roadmap (not yet shipped)

  • Guided mitigation — extending from investigation into helping engineers take the right corrective action safely.
  • Cross-incident learning — using patterns from past incidents to improve future diagnoses, surface recurring issues before they page, and identify systemic reliability gaps.

Comparison to Instacart Blueberry

Both AI SRE and Instacart Blueberry compress the "noisy window right after a page fires" by gathering evidence in parallel and producing grounded, evidence-backed explanations rather than generic suggestions. Distinctive to AI SRE: an explicit four-layer platform that teams extend with their own agentic runbooks, and an explicit emphasis on API-layer guardrails for agent access to shared observability infra.

Seen in

Last updated · 766 distilled / 2,225 read