Skip to content

DATABRICKS 2026-09-24

Read original ↗

How I built agent-based security reviews on Databricks

Summary

A Databricks security leader describes building an agent-based layer on top of existing security-review automation to reclaim expert reviewer time. The problem was queue triage: a routine, familiar-pattern integration sat next to a genuinely novel, high-risk architecture, both waiting on the same scarce experienced reviewer. The solution was not one general-purpose "security reviewer" agent but a set of seven focused agents, each with a bounded responsibility (intake, risk assessment, requirements, specialized review, validation, workflow, learning), orchestrated as scheduled jobs behind a conversational intake application. The whole system was built on the Databricks platform: Unity Catalog as the governed system of record for standards / requests / evidence / decisions, Databricks-hosted foundation models (Claude Haiku / Sonnet / Opus) as the reasoning layer, Lakeflow Jobs as the orchestrator on serverless compute, and Databricks Apps for both the intake experience and an executive dashboard. The load-bearing thesis is about trustworthiness, not model quality: automated completion is confined to predefined low-risk request classes, every decision must be backed by concrete evidence, and any missing/contradictory evidence triggers a conservative escalation to a human rather than an optimistic inference. Author reports a working first version in under two hours (vs. weeks to wire up separate services); the team later expanded it into a shared, hardened production system.

Key takeaways

  1. Decompose into bounded agents, not one authority. The author "deliberately avoided building a single agent with broad authority to act as a security reviewer." Instead, seven agents each own one narrow job so "its behavior remains inspectable and testable, and a change to one does not silently affect another." This is the canonical specialized-agent-decomposition shape applied to security intake/review.
  2. Automate the process, not the judgment. "Automation for well-understood cases within explicit criteria; people for novel, high-risk, or ambiguous decisions." The system automates repeatable steps and preserves human authority where risk or uncertainty is high — high-risk / critical / unusual / ambiguous requests route to a person by default.
  3. Evidence-grounded decisioning with conservative defaults. Three things make an automated decision safe to act on: predefined request classes (only well-understood low-risk categories are eligible), acceptable evidence (concrete, verifiable artifacts — a linked design doc, a stated data-classification level, config demonstrating an approved control; "an assertion with no evidence is treated as missing information"), and a conservative response to uncertainty ("When evidence is missing or contradictory, the system does not guess. It defaults to the more conservative risk tier, posts a specific clarification request, or hands the request to a reviewer... It does not find its way to approval."). This is fail-closed on the security invariant.
  4. Model tiering by task difficulty. "Haiku handles lightweight classification, Sonnet handles most review work, and Opus is reserved for the heaviest reasoning." A cost/capability ladder over the same reasoning layer — the cheap-approximator / expensive-fallback shape and task-difficulty routing.
  5. A better front door beats a static form. A conversational Databricks Apps intake app turns a plain-language description into a structured request, asks context-dependent follow-ups, highlights missing information, and gives a preliminary risk indication before a formal request exists. A consultation mode grounded in the standards lets teams get guidance "while they are still shaping a design" without opening a ticket — "not treating the queue as the only way to engage security."
  6. Grounding is the trust primitive. "The hardest part was not getting a model to produce an answer. It was making that answer constrained, reviewable, and appropriate for action." Agents are grounded in the organization's security standards (RAG-style grounding), and the requirements agent produces "specific, checkable items tied to that architecture" rather than "generic boilerplate" — structured, checkable output.
  7. One governed system of record. Unity Catalog holds standards, request data, evidence, model outputs, and decisions as governed tables under "one permission and lineage model across all of it." Because metrics come from the same operational records, they trace back to the requests and decisions that generated them — an audit trail by construction, and the reason the executive dashboard's numbers are trustworthy without reconciling exports across systems.
  8. A learning agent closes the loop without changing prod behavior. A dedicated learning agent "periodically compares reviewer edits against the original output to surface improvements to prompts and standards." Corrections refine standards / prompts / workflow logic — humans apply the change; the agents do not mutate production behavior on their own. This is the human-correction-as-signal loop.
  9. Platform consolidation was the velocity lever. Building data, models, workflows, and applications in one environment gave "a consistent governance and operational model" — the author had a working system in "under two hours" vs. "weeks" to wire up separate services with different permissions, logs, and data paths.

Architecture

A request flows through one continuous path, every step reading from and writing to the same governed tables:

  1. Intake — a conversational Databricks Apps app turns a plain-language description into a structured request and attaches supporting design documents.
  2. Reasoning — Databricks-hosted foundation models (Claude Haiku / Sonnet / Opus via the Foundation Model API) classify the request, assess risk, and draft requirements, always grounded in the security standards.
  3. Orchestration — Lakeflow Jobs run the review agents on serverless compute, moving each request through its stages on a schedule.
  4. System of record — Unity Catalog holds standards, request data, evidence, model outputs, and decisions as governed tables with one permission + lineage model.
  5. Observability — a second Databricks App reads those same tables to report volume, risk mix, automation rate, and time saved.

The seven focused agents

  • Intake agent — runs the conversational front door: identifies the review path, asks context-dependent questions, assembles a structured request.
  • Risk assessment agent — assigns a risk tier with supporting evidence, and defaults to a higher tier when the picture is incomplete.
  • Requirements agent — maps a request to relevant standards and drafts implementation-specific requirements for well-understood cases.
  • Specialized review agents — handle request types needing dedicated logic (e.g. browser-extension threat modeling, third-party vendor assessment).
  • Validation agent — builds a per-item validation checklist for higher-risk requests before anything closes.
  • Workflow agent — handles follow-ups: clarifications, reminders, acknowledgment tracking, and person escalations.
  • Learning agent — periodically compares reviewer edits against original output to surface prompt / standard improvements.

Worked example (routine case)

An internal integration on an approved SSO pattern handling no sensitive data. Instead of boilerplate, the requirements agent produces specific, checkable items tied to that architecture, e.g.: authenticate through the approved IdP and disable local/shared credentials; restrict to minimum necessary access scopes and document them; send app + access logs to the central logging pipeline; confirm data-classification level and re-review before any sensitive data is introduced. The requester acknowledges each item and attaches evidence; if everything checks out and the case meets eligibility criteria, it completes via the automated path. Anything high-risk / critical / unusual / ambiguous goes to a person, who receives a structured summary, supporting evidence, applicable standards, and open questions.

Operational numbers / claims

  • 7 focused agents in the review set.
  • 3-tier model ladder: Claude Haiku (classification) → Sonnet (most review work) → Opus (heaviest reasoning).
  • Working first version in under 2 hours (vs. "weeks" to wire up separate services).
  • 2 Databricks Apps: conversational intake + executive dashboard.
  • Dashboard tracks: request volume, risk distribution, automated completion, human escalation, cycle time, estimated reviewer time saved. (Specific dashboard figures are redacted in the post.)
  • Reported outcome: "Eligible routine requests that once waited in the queue for days can now be completed in minutes."

Caveats

  • Vendor blog / first-person account; no independent verification and the concrete dashboard numbers are redacted.
  • The post does not spec the inter-agent coordination protocol, the exact eligibility-criteria representation, or how the risk-tier thresholds are encoded.
  • The single specific quantitative outcome ("days → minutes" for eligible routine requests) is illustrative, not a measured aggregate.
  • The whole system runs inside one vendor platform (Databricks); the consolidation-as-velocity claim is inseparable from that lock-in.

Source

Last updated · 766 distilled / 2,225 read