Databricks — Data Ontology defined: The context layer your AI agents are missing¶
Summary¶
An interview with Richard Tomlinson (Databricks) arguing that the missing layer of enterprise AI infrastructure is the data ontology — a semantic context layer that captures what data means in the business (definitions, relationships, calculations, authoritative sources, expertise, permissions), as distinct from a schema, which only captures how data is structured. The thesis: for decades a knowledgeable human analyst sat between ambiguous data and the final report and supplied the missing context as tribal knowledge; AI agents break that assumption because no human is in the loop, so the agent must discover meaning itself or fill the gap with plausible-but-fabricated inference. Past attempts (manual semantic layers, enterprise knowledge graphs) became shelfware because teams tried to model the entire enterprise up front, and the model went stale faster than it could be built. The proposed alternative — "model the head and learn the tail" — is to explicitly govern the small set of concepts that cannot be wrong (revenue, compliance rules, core KPIs) while continuously learning the long tail from existing dashboards, queries, notebooks, documents, and usage, ranking knowledge by authority and respecting governance. Databricks' Genie Ontology is offered as the concrete implementation (the automatic context layer under Genie One / Genie Agents), with a benchmark result and permission-enforcement-at-retrieval design as evidence.
Key takeaways¶
-
Schema ≠ ontology. A schema tells an agent a table has
net_rev,gross_rev,recog_rev; an ontology tells it which definition of revenue applies to the question, which source Finance considers authoritative, how the calculation is normally performed, and whether the asker is even permitted to see it. "The schema is the map of the data. The ontology is closer to a map of how the organization understands and uses that data." (Source: sources/2026-09-15-databricks-data-ontology-defined-the-context-layer-your-ai-agents-are-missing) → concepts/data-ontology -
Agents break the "intelligent-user-supplies-context" assumption. Enterprise data architecture has always assumed a human analyst supplies the missing business context ("Which revenue? For which business unit?"). Agents have no human in the loop, so the context humans carried in their heads — definitions, relationships, calculations, authoritative sources, expertise, permissions — must now be explicit infrastructure.
-
The dangerous failure mode is a fluent, specific, wrong answer. In internal testing, an assistant asked to prepare a Product Advisory Board briefing confidently stated 24 customers were participating; challenged, it admitted it had fabricated the number. "A plausible answer can be more dangerous than no answer at all" — the model doesn't know which internal source holds ground truth, so it fills the gap with inference. → concepts/llm-hallucination
-
Why semantic layers / knowledge graphs became shelfware: they tried to model everything. Business knowledge changes too fast and lives in too many places (a dashboard, a SQL query, a notebook, a ticket, a team's habits) for a central team to document and keep current. If maintaining the ontology becomes a separate enterprise data-modeling project, it falls behind the business it describes.
-
"Model the head and learn the tail." Humans explicitly define and govern the small set of concepts that cannot be wrong (revenue, compliance rules, core KPIs); the broader ontology continuously learns the long tail from existing artifacts and operations, ranking knowledge by authority and respecting governance. You don't build a formal ontology from scratch — most of the business understanding already exists in metric definitions, certified data, dashboards, queries, notebooks, and repeated usage.
-
Better context can matter more than more reasoning time. On an internal Databricks benchmark of 28 real-world enterprise data-analysis questions, Genie with Ontology answered 84.5% correctly on the first attempt, vs 52.4% for the strongest general-purpose coding agent in the same evaluation — and Genie was roughly 2× faster. Stated as an internal benchmark, not a universal guarantee. (Source: same) → systems/databricks-genie-ontology
-
An ontology doesn't make an agent accurate by itself. What improves accuracy is supplying the right, authoritative context at the point where the agent is reasoning — which definition applies, where the trusted data lives, which relationships/calculations matter — so the agent spends less time guessing and exploring wrong paths.
-
Governance has two jobs, and both feed the ontology. (a) Access — the ontology must never become a back door around permissions; with Genie Ontology, permissions are enforced during retrieval using the underlying sources' governance (including Unity Catalog), so two employees asking the same question can appropriately get different answers. (b) Trust — certification, authoritative definitions, lineage, usage, expertise, and source provenance become the signals that teach the agent what to trust (the official revenue definition vs a one-off calculation from six months ago). → concepts/governed-agent-data-access, concepts/data-lineage
-
Governance/analytics assets are now strategic AI assets. Metric definitions, documentation, lineage, certifications, usage patterns, and business glossaries "collectively teach AI how the company works." The emerging architecture is not just a data plane + a model — it needs a shared context layer that serves the same business understanding to many agents and applications.
-
The "context tax" already exists, pre-agents. Analysts spend time rediscovering definitions, locating authoritative sources, and reconciling conflicting reports; teams recreate the same semantics inside different BI tools; self-service breaks the moment a question gets nuanced. Agents make this existing cost visible — without context they repeat the discovery process computationally (exploring schemas, reading docs, trying queries), adding latency, token spend, and cost with no guarantee of correctness.
Systems / concepts / patterns extracted¶
- Systems: Genie Ontology (the automatic context layer under Genie One + Genie Agents), Genie / Genie One / Genie Agents, Unity Catalog (source-of-truth governance; permission enforcement at retrieval), Metric Views (the governed metric/semantic layer that supplies authoritative definitions).
- Concepts: concepts/data-ontology (NEW — semantic context layer vs schema), concepts/llm-hallucination (fabricated "24 customers"), concepts/governed-agent-data-access (permission-at-retrieval; ontology not a back door), concepts/knowledge-graph (ontology as relationships/definitions substrate), concepts/text-to-sql (the NL→query path grounded by the ontology), concepts/data-lineage + concepts/fine-grained-authorization (trust/authority + access signals).
- Design principle (recorded as prose, not minted): "model the head and learn the tail" — explicitly govern the critical-concept head, auto-learn the long tail from existing artifacts + usage while ranking by authority. Strongly attested here and adjacent to the 2026-09-03 Databricks "Governance beyond security: knowledge, context & ontology on the lakehouse" post; kept as a tag/prose pending a clearer second independent source per the taxonomy gate.
Operational numbers¶
- 28 real-world enterprise data-analysis questions in the internal benchmark.
- 84.5% first-attempt correctness for Genie with Ontology vs 52.4% for the strongest general-purpose coding agent in the same eval.
- ~2× faster than that coding agent (Genie).
- Fabricated figure in the failure anecdote: 24 customers (invented, not grounded).
Caveats¶
- Vendor thought-leadership / product-positioning format. Interview promoting Genie Ontology; the architectural claims are conceptual, and the benchmark is an internal Databricks benchmark (28 questions), explicitly "not a universal accuracy guarantee." Baseline coding agent is unnamed.
- Mechanism-light. How the "learn the tail" inference actually works (what signals, how definitions/relationships are inferred, how authority is ranked, how staleness is bounded) is not disclosed here — the piece is a definition + design-philosophy post, not an internals disclosure.
- Tier-3 source, borderline-include: it is a launch-adjacent interview, but the architecture content (ontology-vs-schema, permission-at-retrieval, model-the-head/learn-the-tail, authority-ranking, the benchmark) clears the ">20% architecture" bar per AGENTS.md borderline rule.
Source¶
- Original: https://www.databricks.com/blog/data-ontology-defined-context-layer-your-ai-agents-are-missing
- Raw markdown:
raw/databricks/2026-09-15-data-ontology-defined-the-context-layer-your-ai-agents-are-m-c52ed0ba.md
Related¶
- concepts/data-ontology — the core concept this post defines.
- systems/databricks-genie-ontology — the implementation.
- systems/databricks-genie — the analytics agent Genie Ontology grounds.
- concepts/llm-hallucination — the failure mode context prevents.
- concepts/governed-agent-data-access — permission enforcement at retrieval; ontology not a back door around ACLs.
- concepts/knowledge-graph — adjacent framing (relationships/definitions substrate); see Netflix UDA and Zalando MDM instances.
- systems/unity-catalog · systems/databricks-metric-views — the governance + governed-semantic-layer substrate.
- sources/2026-09-03-databricks-governance-beyond-security-knowledge-context-ontology-on-the-lakehouse — companion Databricks post; governance-as-knowledge/context/ontology.