Skip to content

DATABRICKS 2026-09-23

Read original ↗

How Concurrence governs clinical AI at a trillion-token scale with Unity Gateway

Summary

A Databricks customer story documenting how Concurrence — a healthcare company building clinical AI agents (AI clinicians, nurses, care coordinators, ambient documentation, care-plan summaries, knowledge retrieval) — runs production agentic systems at roughly 1.2 trillion annualized input tokens. The load-bearing architectural ideas are: (1) an immutable-event "world model" where new patient information is appended as events and current patient state is computed from history (event sourcing), giving agents consistent, provenance-preserving context and enabling replay-based testing; (2) a shared Databricks data foundation — Zerobus Ingest → governed Delta tables, Spark Declarative Pipelines deriving the world model, Unity Catalog per-tenant governance, Lakebase serving operational state; and (3) compliance-gated model routing through Unity Gateway — models must clear compliance (HIPAA / BAA coverage) before quality, performance, or cost are considered. Developer AI (coding agents) already runs fully through the Gateway's ug CLI; batch inference runs on BAA-covered Databricks-hosted Claude; real-time patient/clinician inference is built and feature-flagged behind a synthetic canary, pending compliance coverage.

Key takeaways

  1. Scale. ~100.8B input tokens and 11.2M LLM calls per 30 days (~1.2T annualized input tokens); monthly token volume grew ~5x from a late-2025 baseline to July 2026. (Source: this article.)
  2. World model = event sourcing for patient state. Because healthcare data conflicts across systems and the newest value is not always the most reliable, Concurrence records new information as immutable events rather than overwriting records, then computes current patient state from that history while preserving source and provenance. This is a textbook instance of log-as-truth / database-as-cache applied to a clinical world model. What agents learn from patients/clinicians feeds back into the state for future workflows.
  3. The event history makes replay-based agent testing possible. Because patient state is derived from an immutable event log, teams can replay that state and test different paths without mutating the real patient record — evaluating how an agent responds to scenarios before it reaches a real patient (patterns/snapshot-replay-agent-evaluation). Simulation + evaluation traffic is ~7x production traffic on the new platform.
  4. Data foundation. Events stream through Zerobus Ingest into governed Delta tables — 2.7M world-model events/month, 90k/day at peak. Spark Declarative Pipelines derive the world model + clinical data; Unity Catalog governs each customer environment; Lakebase serves patient state, agent + conversation state, and knowledge-base data to operational apps. The move retired a homegrown prompt-log store and reverse-ETL jobs in favor of governed Delta tables + Lakebase Synced Tables.
  5. Compliance-gated routing (the headline design principle). "Routing is compliance-gated before it is cost-gated. Models must first meet the compliance requirements of a workload before Concurrence considers quality, performance or cost." For batch AI, ai_query jobs run on Databricks-hosted Claude under a BAA, and the endpoint resolver only permits models within the BAA-covered namespace, structurally preventing PHI from being routed to an uncovered model. The same covered path also runs safety classification (self-harm, suicidal ideation, medical emergencies) — the highest-stakes workloads protected by the same architectural constraint.
  6. Per-tenant isolation via Unity Catalog. Each healthcare organization gets its own schema and service principal, with access controls, lineage, and audit trails governed through Unity Catalog (concepts/tenant-isolation, concepts/audit-trail). HIPAA-, GDPR-, SOC 2-compliant today; HITRUST and ISO 27001/42001 in progress.
  7. Coding-agent governance at scale. All coding-agent model + tool traffic routes through Unity Gateway's coding CLI, ug. Each request stays associated with the identity of the person who made it (patterns/on-behalf-of-agent-authorization); MCP access is centrally managed with permissions by engineer group and per-user auth to tools like Databricks, Datadog, and Linear. In July, 14 users generated 35.85B input tokens, 95.37% of them cache reads; since ug rolled out July 10, ~360,000 requests and ~61B cumulative input tokens.
  8. Attribution + cost visibility. Every request is attributed to the engineer (ug usage for individuals); org-level tracking uses system.ai_gateway.usage for models-in-use, token consumption, cache rates, and spend by person/team (patterns/telemetry-to-lakehouse).
  9. Multi-model, governed catalog. 14 models serving production inference; a governed catalog of 46 models. Most production volume runs on smaller, faster models with frontier models reserved for complex reasoning. Claude Opus 4.8 and GPT-5.6 Sol account for most coding-agent usage, Opus 5 growing. Concurrence is developing clinical reasoning benchmarks and testing Unity Gateway Smart Routing against its own healthcare-specific routing — where model eligibility starts with compliance, so routing optimizes only within each workload's compliance boundary. Exploring Omnigent as a coding-agent meta-harness.
  10. Compounding foundation. Each agent's work enriches the patient state the next agent starts from, so new workflows reuse existing context rather than rebuild it — reducing the incremental cost of adding AI workflows (patterns/async-projected-read-model over the world-model log).

Systems

  • systems/concurrence — the healthcare clinical-AI platform (new page).
  • Unity Gateway — centralized inference + tool control point; coding CLI ug; compliance-gated routing; Smart Routing tests.
  • Unity Catalog — per-tenant schema + service principal, lineage, audit; system.ai_gateway.usage billing view.
  • Lakebase — operational serving of patient / agent / conversation / knowledge-base state; Synced Tables replacing reverse-ETL.
  • Zerobus Ingest — event + agent-trace ingest to Delta.
  • Spark Declarative Pipelines — derive the world model + clinical data from the event log.
  • Databricks Apps — care-packet guide, nurse care-plan summary, clinical-content review surface.
  • Foundation Model API — BAA-covered Databricks-hosted Claude for ai_query batch scoring.
  • MCP — governed tool surface (Databricks, Datadog, Linear) via ug.
  • Omnigent — meta-harness under evaluation.

Concepts / patterns

Operational numbers

Metric Value
Input tokens / 30 days ~100.8B
LLM calls / 30 days 11.2M
Annualized input tokens ~1.2T
Monthly token growth (late-2025 → Jul 2026) ~5x
World-model events / month 2.7M (90k/day peak)
Simulation+eval vs production traffic ~7x
Coding-agent users (July) 14
Coding-agent input tokens (July) 35.85B
Cache-read fraction (July) 95.37%
Coding-agent requests since Jul 10 ~360,000
Cumulative coding-agent input tokens since Jul 10 ~61B
Production inference models 14
Governed model catalog 46

Caveats

  • Customer story, not an internals post. Databricks-authored; light on mechanism (no world-model schema, no SDP DAG, no resolver/router internals, no Lakebase sizing, no p99 latency).
  • "Compliance-gated routing" is disclosed as a principle, not a mechanism — the only concrete lever named is the endpoint resolver restricting to the BAA-covered namespace.
  • Real-time patient/clinician inference is not yet on Unity Gateway — it's built, feature-flagged, and canary-tested, pending compliance coverage; today it remains on Concurrence's existing provider infrastructure.
  • Tier-3 source (Databricks); ingested for the event-sourced world model, replay-based evaluation, compliance-gated routing design, and production-scale numbers.

Source

Last updated · 766 distilled / 2,225 read