Skip to content

SYSTEM Cited by 1 source

Concurrence

Concurrence is a healthcare company operating clinical AI agents across patient- and provider-facing workflows — AI clinicians, nurses, and care coordinators; ambient documentation; care-plan summaries; and knowledge retrieval. It runs these agentic systems in production at roughly 1.2 trillion annualized input tokens (~100.8B input tokens and 11.2M LLM calls per 30 days, ~5x monthly growth from a late-2025 baseline to July 2026), consolidating its data, governance, and AI-access stack on Databricks. (Source: sources/2026-09-23-databricks-how-concurrence-governs-clinical-ai-at-a-trillion-token-scal)

Concurrence is a proper-noun customer system (not a Databricks product); it is documented here because the article discloses substantive architecture — an event-sourced clinical "world model," replay-based agent evaluation, and compliance-gated model routing — at production scale.

The world model — event sourcing for patient state

Healthcare data conflicts across systems, and the newest value is not always the most reliable. Rather than overwriting records, Concurrence records new information as immutable events and computes the current patient state from that history, preserving the source and provenance of each fact. This derived, provenance-preserving state is what it calls its world model. Agents get a consistent view and history of what is known about a patient, and what they learn from patients and clinicians feeds back into the state for future workflows.

This is a clinical instance of the log is truth, the database is a cache: the immutable event log is authoritative; patient state is a materialized, replayable projection of it (patterns/async-projected-read-model).

Care-gap and medication-adherence outreach is the worked example: agents call or text patients overdue for follow-up care or falling off a medication, use existing patient context to guide the conversation, and record what they learn back into the patient state.

Data foundation on Databricks

events (patient info, agent traces)
        │
        ▼
  Zerobus Ingest ──► governed Delta tables   (2.7M world-model events/mo, 90k/day peak)
        │
        ▼
  Spark Declarative Pipelines ──► world model + clinical data
        │                         (Unity Catalog governs each tenant)
        ▼
  Lakebase ──► serves patient state, agent + conversation state, knowledge base
        │
        ▼
  operational apps (Databricks Apps) + agents
  • Zerobus Ingest streams events (and agent traces) directly into governed Delta tables.
  • Spark Declarative Pipelines derive the world model and clinical data.
  • Unity Catalog governs each customer environment — each healthcare org gets its own schema and service principal, with access controls, lineage, and audit trails (concepts/tenant-isolation, concepts/audit-trail).
  • Lakebase serves the operational patient state, agent + conversation state, and knowledge-base data. The migration retired a homegrown prompt-log store and reverse-ETL jobs in favor of governed Delta tables + Lakebase Synced Tables.
  • Production apps run on Databricks Apps — a care-packet guide, a nurse care-plan summary, and a clinical-content review surface, all on the same governed patient context.

Testing agents before they touch a patient

Because patient state is derived from an immutable event history, teams can replay that state and test different paths without mutating the real patient record, evaluating how an agent responds to scenarios before deploying it (patterns/snapshot-replay-agent-evaluation). On the new platform, simulation + evaluation traffic is ~7x production traffic.

Agent traces stream through Zerobus Ingest and land in Delta alongside the clinical data that produced them. Scheduled ai_query jobs (Databricks-hosted Claude) score interactions for conversation quality and safety, extract memory, and write results back to Delta — so patient data, traces, outcomes, and evaluations sit on one governed foundation and teams can attribute a performance change to the model, the data, or the workflow (patterns/telemetry-to-lakehouse).

Compliance-gated model routing

HIPAA shapes the architecture from the start (Concurrence is HIPAA-, GDPR-, and SOC 2-compliant; HITRUST and ISO 27001/42001 in progress). The defining routing principle: routing is compliance-gated before it is cost-gated — a model must first meet a workload's compliance requirements before quality, performance, or cost are considered. See concepts/model-first-routing for the generalized form and concepts/centralized-ai-governance for the governance framing.

  • Batch AI: ai_query jobs run on Databricks-hosted Claude under a BAA. The endpoint resolver only permits models within the BAA-covered namespace, structurally preventing PHI from being routed to an uncovered model. The same covered path runs safety classification for self-harm, suicidal ideation, and medical emergencies — highest-stakes workloads protected by the same architectural constraint.
  • Real-time patient/clinician inference remains on Concurrence's existing provider infrastructure today. The Unity Gateway integration is built and feature-flagged, with a synthetic canary continuously testing it end to end; production traffic can move as compliance coverage becomes available.
  • Multi-model: 14 models in production inference, a governed catalog of 46 models; most volume on smaller/faster models, frontier models reserved for complex reasoning. Concurrence is building clinical reasoning benchmarks and testing Unity Gateway Smart Routing against its own healthcare-specific routing — optimizing model choice within each workload's compliance boundary.

Governing coding agents with ug

All coding-agent model + tool traffic routes through Unity Gateway's coding CLI, ug. Developers get one governed path to approved models + MCP tools, with each request tied to the identity of the person who made it (patterns/on-behalf-of-agent-authorization). MCP access is centrally managed, permissions assigned by engineer group, and each user authenticates individually to tools such as Databricks, Datadog, and Linear. This includes forward-deployed engineers working inside customer environments that handle sensitive healthcare data.

  • July: 14 users, 35.85B input tokens, 95.37% cache reads (concepts/cache-hit-rate).
  • Since ug rolled out July 10: ~360,000 requests, ~61B cumulative input tokens.
  • Attribution: individuals monitor usage via ug usage; org-level tracking via system.ai_gateway.usage (models, tokens, cache rates, spend by person/team).
  • Claude Opus 4.8 and GPT-5.6 Sol dominate coding-agent usage; Opus 5 growing. Exploring Omnigent as a coding-agent meta-harness.

What the article does not disclose

  • World-model event schema; SDP pipeline DAG; state-computation logic.
  • Endpoint-resolver / router internals beyond "restrict to BAA-covered namespace."
  • Lakebase sizing, latency, and the real-time serving path's shape.
  • Per-workflow p50/p99 latency; cost-per-token; safety-classifier model.
Last updated · 766 distilled / 2,225 read