Databricks — Governance beyond security: knowledge, context & ontology on the lakehouse¶
Summary¶
A Databricks Data Empowerment Program (DEP) post arguing that data governance for AI is not just security ("who can touch data") but knowledge, context, and ontology ("what the data means, whether it can be trusted, whether an AI should learn from it"). Its durable system-design content is the catalog-centered agentic architecture: governance artifacts (classification tags, comments, certified flags, lineage, glossary terms, de-identification policies, model cards, data contracts) live in Unity Catalog as structured, machine-readable metadata, and AI agents execute against that metadata as their runtime — "build agents" that assemble/deliver data products and "analytic agents" that answer questions on top of them. Each agent runs a read-instructions → do-one-task → write-evidence-back loop. Access is enforced at the data layer via ABAC, not the application layer, so RAG/vector retrieval inherits the querying user's row/column entitlements; anything the catalog doesn't describe is suppressed (fail-closed). The economic thesis: govern the data well and a cheaper/smaller model delivers trustworthy results because the meaning lives in the catalog, not the token bill.
Key takeaways¶
- Governance is reframed as semantics, not just security. "Security tells you who can touch data. It says nothing about what the data means, whether it can be trusted, or whether an AI model should ever learn from it." Every governance artifact is recast as semantic raw material — "Every classification tag is a concept. Every model card is context. Every data contract is a shared definition. Every lineage link is a relationship." (Source: sources/2026-09-03-databricks-governance-beyond-security-knowledge-context-ontology-on-the-lakehouse) See concepts/centralized-ai-governance.
- The catalog is the agent runtime, not a doc store. "The catalog isn't where you just document governance; it's the runtime the agents execute against." Everything a build agent needs — source-to-target mappings, business definitions, classification tiers, de-identification policies, data contracts, model cards — lives in Unity Catalog as governed metadata. See catalog-as-agent-runtime.
- Agents run a read → act → write-back loop against the catalog. "It reads instructions from the catalog; does one concrete task … and then writes the evidence back as test outcomes, quality scores, lineage, or change capture data. This repeats." Five build agents (ETL generation, testing, curation validation, de-identification, deployment) consume catalog metadata and write results back to it. See catalog-metadata-as-agent-instruction-set.
- Two agent kinds, each bound to a single data product. "Build agents" assemble/deliver data products; "analytic agents" (e.g. a Genie Agent) answer business questions — "each bound to a single data product." Binding one agent to one domain-scoped certified product is "the single largest accuracy lever available" and the accountability mechanism: a wrong answer from a misdefined metric goes to the Data Product Owner, not the AI team. See agent-bound-to-single-certified-data-product.
- Access is enforced at the data layer via ABAC, spanning SQL and vector search alike. "Access controls operate via Attribute-Based Access Control (ABAC) at the data layer, not the application layer. If a user cannot query a row in SQL, no agent can retrieve it via vector search or embeddings." Agents act with the querying user's entitlements, not a privileged service account. See unified-permission-model-over-data-and-embeddings.
- Fail-closed by default: suppress the undescribed. "If the catalog doesn't explicitly describe a data asset, the system defaults to suppression rather than guessing." Analytic agents decline to answer ("This dataset lacks active certification or semantic mapping required to process your request."); build agents halt before staging and log an unmapped-asset flag for steward review. See patterns/default-on-security-upgrade.
- De-identification is driven by security-policy metadata, not spreadsheets. A three-step Discover → Curate → Execute flow: discovery scanners + InfoSec policy engines classify sensitive columns; classifications land in the catalog as curated policy metadata; the de-id agent reads that curation and produces synthetic or de-identified, HIPAA Safe Harbor-compliant, referentially intact data. "InfoSec policies stop being PDFs and become executable." See de-identification-from-security-policy-metadata.
- AI Certification is a queryable scorecard with continuous expiration. Recorded in Unity Catalog across four dimensions (Governance, Quality, Semantics auto-computed from system tables; Ownership requires a steward signature). "A schema change, contract update, or failed evaluation suite instantly revokes certification until checks … rerun and pass." Non-prod (SIT/regression/model-testing) consumes synthetic or de-identified data only — production PHI never leaves the governed boundary.
- Auto-proving + continuously-improving-context lifecycle. Two tracks (data product + the agent/model on top) move through the same five gates to one shared certification. The lifecycle is "auto-proving" (proof of trustworthiness is a byproduct of delivery, not an audit fire-drill) and feeds an AgentOps loop — "failed queries, hallucination clusters, and user downvotes into the next sprint['s] semantic backlog." Most gates clear in hours inside standard developer tooling; a formal gate meeting is the exception.
- Chase model economics, not model headlines. "When the catalog already supplies the meaning, quality, and context, the model doesn't have to." Frontier models often "mask underlying metadata gaps"; with schemas/business rules explicitly cataloged, "smaller domain-specific models deliver identical accuracy at a fraction of the token cost." See data-readiness-vs-model-spend.
Governance's five pillars (the DEP lens)¶
One lens — every governance artifact contributes to semantics — viewed from five fronts:
- Data Governance — catalog, quality, curation, lineage, with security/compliance built in (PII classification, access control, HIPAA/GDPR, AI-privacy risks).
- Knowledge (AI/ML) Governance — model documentation, responsible-AI standards (bias/fairness, explainability, human oversight, EU AI Act).
- Data Literacy — training, self-service enablement, certification, ROI KPIs.
- Data Management — architecture, data engineering, data-product contracts (schema agreements, SLAs, producer/consumer obligations).
- Ontology — glossary, taxonomy, knowledge graph → an AI semantic layer (LLM context, RAG grounding, chat-query readiness).
The four accountable roles¶
Replacing "vague governance committees":
- Data Product Owner — accountable for a product's definitions/quality; single point of contact when an answer is wrong.
- Data & AI Governance Engineer — translates policy into executable catalog metadata (classifications, contracts, lineage) so rules run at runtime instead of "sitting in a PDF."
- Steward — reviews automated findings, signs release gates ("automation proposes; the steward decides").
- Security / IAM — owns classification tiers and access attributes that drive de-identification and row-level entitlements.
Take-action (the incremental adoption path)¶
Prove the model on one data product before an enterprise overhaul: Scan (enable discovery on one schema) → Define (set certification thresholds in Unity Catalog for completeness/semantics/quality) → Bind (attach one analytic agent + eval suite + de-identified testing path) → Assign (one named Data Product Owner). Then repeat one certified product at a time.
Extracted systems / concepts / patterns¶
- Concepts: catalog-as-agent-runtime (new), concepts/centralized-ai-governance (new), data-readiness-vs-model-spend (new), unified-permission-model-over-data-and-embeddings (new), concepts/centralized-ai-governance, concepts/governed-agent-data-access, concepts/attribute-based-access-control, concepts/fail-open-vs-fail-closed, concepts/governed-agent-data-access, run-with-end-user-credentials, data-product-certification, concepts/knowledge-graph.
- Patterns: catalog-metadata-as-agent-instruction-set (new), de-identification-from-security-policy-metadata (new), agent-bound-to-single-certified-data-product (new), patterns/default-on-security-upgrade.
- Systems: systems/unity-catalog, systems/unity-catalog-abac, systems/unity-catalog-data-classification, systems/databricks-genie, systems/databricks.
Caveats¶
- Tier-3 advocacy/DEP post, healthcare-framed. The catalog-as-runtime architecture, the read→act→write-back agent loop, the data-layer ABAC guarantee spanning vector search, fail-closed suppression, and the de-id-from-classification pipeline are the durable, portable ideas; the five-pillars taxonomy and "Don't chase model headlines" section are framework/opinion content.
- No throughput/latency numbers, no architecture diagrams beyond marketing figures, no code. The "cheaper model at identical accuracy" claim is asserted, not benchmarked in the post.
- Borderline scope call (included): architecture content (catalog runtime, ABAC-over-embeddings, fail-closed, de-id pipeline, certification-as- queryable-scorecard) is ≳50% of the body, above the AGENTS.md ≥20% borderline bar for Tier-3.
- Adjacent to but distinct from the same-day "Building High-Quality and Trusted Data Products" post: that one covers the data-product lifecycle
- contract; this one covers the agentic runtime the governed products feed and the model-economics argument.
Source¶
- Original: https://www.databricks.com/blog/governance-beyond-security-knowledge-context-ontology-lakehouse
- Raw markdown:
raw/databricks/2026-09-03-governance-beyond-security-knowledge-context-ontology-on-the-d4fa99e9.md
Related¶
- catalog-as-agent-runtime — the catalog is the runtime agents execute against.
- concepts/centralized-ai-governance — the reframe: governance builds AI, it isn't a gate in front of it.
- concepts/centralized-ai-governance — the Pillar-2 sibling principle this post extends into an agent runtime.
- catalog-metadata-as-agent-instruction-set — the read→act→write-back agent loop.
- de-identification-from-security-policy-metadata — Discover→Curate→Execute de-id.