Skip to content

CONCEPT Cited by 2 sources

Data ontology

Definition

A data ontology is a semantic context layer that captures what data means in the context of a business — its definitions, relationships, calculations, authoritative sources, expertise, and access rules — as opposed to a schema, which only captures how data is structured (Source: sources/2026-09-15-databricks-data-ontology-defined-the-context-layer-your-ai-agents-are-missing).

"A schema describes how data is structured. An ontology describes what that data means in the context of the business… The schema is the map of the data. The ontology is closer to a map of how the organization understands and uses that data."

Concretely, a schema tells a query engine that a table contains net_rev, gross_rev, and recog_rev. An ontology tells an agent which definition of revenue applies to this question, which source Finance considers authoritative, how the calculation is normally performed, and whether the person asking is even permitted to see it.

Why AI agents make this urgent

For decades enterprise data architecture rested on an implicit assumption: organize the data correctly and an intelligent human user will figure out what it means. Tables, schemas, catalogs, and dashboards provide structure; a knowledgeable analyst supplies the missing context as tribal knowledge — knowing which of five revenue tables Finance trusts, what "active customer" means this quarter, or which dashboard is deprecated.

AI agents break this assumption because there may be no knowledgeable human between the data and the decision. The agent must discover the meaning itself. Giving it more data does not help; it needs the business context humans historically carried in their heads. When that context is missing, the agent fills the gap with plausible inference — the fluent-but-fabricated failure mode (in the source's anecdote, an assistant confidently invented that "24 customers" were attending a board meeting). "In enterprise AI, a plausible answer can be more dangerous than no answer at all."

Ontology vs schema vs knowledge graph

  • Schema — structure only (columns, types, keys).
  • Ontology — meaning: definitions, relationships, calculations, authoritative sources, expertise, permissions. Often realized as a knowledge graph (nodes = business concepts / metrics / assets; edges = mappings + relationships), but the ontology is the semantic contract, not the storage substrate.
  • A schema registry governs message/table shape; an ontology governs interpretation. The two are complementary — Netflix's UDA explicitly "unified a data catalog with a schema registry, but with a hard requirement for semantic integration," which "naturally led us to consider a knowledge graph approach" (see concepts/knowledge-graph).

The shelfware failure mode

Manual semantic layers and enterprise knowledge graphs promised "one version of the truth" for years and mostly became shelfware. The stated cause is not that the idea is wrong but that organizations tried to model the entire enterprise manually. Business knowledge changes too fast and lives in too many places — a dashboard, a SQL query, a notebook, a ticket, or simply how a team repeatedly works — for a central team to document and keep current. Once maintaining the ontology becomes a separate enterprise data-modeling project, it falls behind the business it is supposed to describe.

Model the head, learn the tail

The scalable alternative offered by the source is a design principle: "model the head and learn the tail."

  • Head (govern explicitly). Humans define and govern the small set of concepts that cannot be wrong — revenue, compliance rules, core KPIs.
  • Tail (learn continuously). The broader ontology learns the long tail from existing artifacts and operations — dashboards, queries, notebooks, documents, certified data, and repeated usage — while ranking knowledge by authority and respecting governance.

You do not build a formal ontology from scratch; requiring that recreates the shelfware scalability problem. Most business understanding already exists in metric definitions, certified data, dashboards, queries, notebooks, and documentation — the goal is to preserve human control over the critical concepts while automatically learning the rest. (This principle is recorded as prose rather than a separate pattern page pending a clearer second independent source.)

Accuracy comes from context at the point of reasoning

An ontology "by itself doesn't magically make an agent accurate." What improves accuracy is supplying the right, authoritative context at the point where the agent is reasoning — which definition applies, where the trusted data lives, which relationships/calculations matter — so the agent spends less time guessing and exploring wrong paths.

Evidence cited: on an internal Databricks benchmark of 28 real-world enterprise data-analysis questions, Genie with Ontology answered 84.5% correctly on the first attempt vs 52.4% for the strongest general-purpose coding agent in the same evaluation, at roughly 2× the speed. Presented as an internal benchmark, not a universal guarantee — but it illustrates the core principle: better enterprise context can matter as much as, or more than, giving the model more time to reason.

Governance is part of the ontology

Governance has two jobs, and both feed the context layer:

  1. Access — never a back door. The ontology must not become a way around existing permissions. If a user can't access the source, the agent shouldn't retrieve context derived from it. With Genie Ontology, permissions are enforced during retrieval using the underlying sources' governance, including Unity Catalog — two employees can ask the same question and appropriately get different answers (concepts/governed-agent-data-access).
  2. Trust — teach the agent what to believe. Certification, authoritative definitions, lineage, usage, expertise, and source provenance become signals that distinguish the official revenue definition from a one-off calculation someone made six months ago. "In the AI era, governance is no longer only about controlling data. It's increasingly part of the mechanism that teaches AI which business knowledge deserves authority."

An implication: assets previously viewed as governance/analytics overhead — metric definitions, documentation, lineage, certifications, usage patterns, business glossaries — become strategic AI assets that "collectively teach AI how the company works." The emerging architecture is not just a data plane plus a model; it needs a shared context layer serving the same business understanding to many agents and applications.

Seen in

Merged aliases

  • semantic-layer
  • context-layer
  • business-context-layer
Last updated · 766 distilled / 2,225 read