How Concurrence governs clinical AI at a trillion-token scale with Unity Gateway¶
Summary¶
A Databricks customer story documenting how Concurrence — a healthcare
company building clinical AI agents (AI clinicians, nurses, care coordinators,
ambient documentation, care-plan summaries, knowledge retrieval) — runs
production agentic systems at roughly 1.2 trillion annualized input tokens.
The load-bearing architectural ideas are: (1) an immutable-event "world
model" where new patient information is appended as events and current patient
state is computed from history (event sourcing), giving agents consistent,
provenance-preserving context and enabling replay-based testing; (2) a shared
Databricks data foundation — Zerobus Ingest →
governed Delta tables, Spark Declarative Pipelines
deriving the world model, Unity Catalog per-tenant
governance, Lakebase serving operational state; and (3)
compliance-gated model routing through Unity
Gateway — models must clear compliance (HIPAA / BAA coverage) before quality,
performance, or cost are considered. Developer AI (coding agents) already runs
fully through the Gateway's ug CLI; batch inference runs on BAA-covered
Databricks-hosted Claude; real-time patient/clinician inference is built and
feature-flagged behind a synthetic canary, pending compliance coverage.
Key takeaways¶
- Scale. ~100.8B input tokens and 11.2M LLM calls per 30 days (~1.2T annualized input tokens); monthly token volume grew ~5x from a late-2025 baseline to July 2026. (Source: this article.)
- World model = event sourcing for patient state. Because healthcare data conflicts across systems and the newest value is not always the most reliable, Concurrence records new information as immutable events rather than overwriting records, then computes current patient state from that history while preserving source and provenance. This is a textbook instance of log-as-truth / database-as-cache applied to a clinical world model. What agents learn from patients/clinicians feeds back into the state for future workflows.
- The event history makes replay-based agent testing possible. Because patient state is derived from an immutable event log, teams can replay that state and test different paths without mutating the real patient record — evaluating how an agent responds to scenarios before it reaches a real patient (patterns/snapshot-replay-agent-evaluation). Simulation + evaluation traffic is ~7x production traffic on the new platform.
- Data foundation. Events stream through Zerobus Ingest into governed Delta tables — 2.7M world-model events/month, 90k/day at peak. Spark Declarative Pipelines derive the world model + clinical data; Unity Catalog governs each customer environment; Lakebase serves patient state, agent + conversation state, and knowledge-base data to operational apps. The move retired a homegrown prompt-log store and reverse-ETL jobs in favor of governed Delta tables + Lakebase Synced Tables.
- Compliance-gated routing (the headline design principle). "Routing is
compliance-gated before it is cost-gated. Models must first meet the
compliance requirements of a workload before Concurrence considers quality,
performance or cost." For batch AI,
ai_queryjobs run on Databricks-hosted Claude under a BAA, and the endpoint resolver only permits models within the BAA-covered namespace, structurally preventing PHI from being routed to an uncovered model. The same covered path also runs safety classification (self-harm, suicidal ideation, medical emergencies) — the highest-stakes workloads protected by the same architectural constraint. - Per-tenant isolation via Unity Catalog. Each healthcare organization gets its own schema and service principal, with access controls, lineage, and audit trails governed through Unity Catalog (concepts/tenant-isolation, concepts/audit-trail). HIPAA-, GDPR-, SOC 2-compliant today; HITRUST and ISO 27001/42001 in progress.
- Coding-agent governance at scale. All coding-agent model + tool traffic
routes through Unity Gateway's coding CLI,
ug. Each request stays associated with the identity of the person who made it (patterns/on-behalf-of-agent-authorization); MCP access is centrally managed with permissions by engineer group and per-user auth to tools like Databricks, Datadog, and Linear. In July, 14 users generated 35.85B input tokens, 95.37% of them cache reads; sinceugrolled out July 10, ~360,000 requests and ~61B cumulative input tokens. - Attribution + cost visibility. Every request is attributed to the engineer
(
ug usagefor individuals); org-level tracking usessystem.ai_gateway.usagefor models-in-use, token consumption, cache rates, and spend by person/team (patterns/telemetry-to-lakehouse). - Multi-model, governed catalog. 14 models serving production inference; a governed catalog of 46 models. Most production volume runs on smaller, faster models with frontier models reserved for complex reasoning. Claude Opus 4.8 and GPT-5.6 Sol account for most coding-agent usage, Opus 5 growing. Concurrence is developing clinical reasoning benchmarks and testing Unity Gateway Smart Routing against its own healthcare-specific routing — where model eligibility starts with compliance, so routing optimizes only within each workload's compliance boundary. Exploring Omnigent as a coding-agent meta-harness.
- Compounding foundation. Each agent's work enriches the patient state the next agent starts from, so new workflows reuse existing context rather than rebuild it — reducing the incremental cost of adding AI workflows (patterns/async-projected-read-model over the world-model log).
Systems¶
- systems/concurrence — the healthcare clinical-AI platform (new page).
- Unity Gateway — centralized inference + tool
control point; coding CLI
ug; compliance-gated routing; Smart Routing tests. - Unity Catalog — per-tenant schema + service principal,
lineage, audit;
system.ai_gateway.usagebilling view. - Lakebase — operational serving of patient / agent / conversation / knowledge-base state; Synced Tables replacing reverse-ETL.
- Zerobus Ingest — event + agent-trace ingest to Delta.
- Spark Declarative Pipelines — derive the world model + clinical data from the event log.
- Databricks Apps — care-packet guide, nurse care-plan summary, clinical-content review surface.
- Foundation Model API — BAA-covered
Databricks-hosted Claude for
ai_querybatch scoring. - MCP — governed tool surface (Databricks,
Datadog, Linear) via
ug. - Omnigent — meta-harness under evaluation.
Concepts / patterns¶
- concepts/log-as-truth-database-as-cache — the world model is event sourcing: immutable events are truth, patient state is a derived cache.
- concepts/centralized-ai-governance — production + batch + developer AI under one inference control point.
- concepts/model-first-routing — compliance-gated variant: eligibility is gated on compliance before capability/cost.
- concepts/tenant-isolation / concepts/audit-trail / concepts/data-lineage — per-org schema + service principal, UC-governed audit + lineage.
- concepts/cache-hit-rate — 95.37% cache reads dominate coding-agent token economics.
- patterns/snapshot-replay-agent-evaluation — replay derived state to test agents without mutating the record.
- patterns/telemetry-to-lakehouse — traces + usage land in governed Delta;
ai_queryscores them for quality/safety and writes results back. - patterns/ai-gateway-provider-abstraction / patterns/central-proxy-choke-point — the Gateway as the single governed path.
- patterns/on-behalf-of-agent-authorization — per-request identity for coding agents and MCP tools.
Operational numbers¶
| Metric | Value |
|---|---|
| Input tokens / 30 days | ~100.8B |
| LLM calls / 30 days | 11.2M |
| Annualized input tokens | ~1.2T |
| Monthly token growth (late-2025 → Jul 2026) | ~5x |
| World-model events / month | 2.7M (90k/day peak) |
| Simulation+eval vs production traffic | ~7x |
| Coding-agent users (July) | 14 |
| Coding-agent input tokens (July) | 35.85B |
| Cache-read fraction (July) | 95.37% |
| Coding-agent requests since Jul 10 | ~360,000 |
| Cumulative coding-agent input tokens since Jul 10 | ~61B |
| Production inference models | 14 |
| Governed model catalog | 46 |
Caveats¶
- Customer story, not an internals post. Databricks-authored; light on mechanism (no world-model schema, no SDP DAG, no resolver/router internals, no Lakebase sizing, no p99 latency).
- "Compliance-gated routing" is disclosed as a principle, not a mechanism — the only concrete lever named is the endpoint resolver restricting to the BAA-covered namespace.
- Real-time patient/clinician inference is not yet on Unity Gateway — it's built, feature-flagged, and canary-tested, pending compliance coverage; today it remains on Concurrence's existing provider infrastructure.
- Tier-3 source (Databricks); ingested for the event-sourced world model, replay-based evaluation, compliance-gated routing design, and production-scale numbers.
Source¶
- Original: https://www.databricks.com/blog/how-concurrence-governs-clinical-ai-trillion-token-scale-unity-gateway
- Raw markdown:
raw/databricks/2026-09-23-how-concurrence-governs-clinical-ai-at-a-trillion-token-scal-839a215b.md