Skip to content

SYSTEM Cited by 2 sources

Databricks Genie Ontology

Genie Ontology is Databricks' automatic context layer underneath Genie One and Genie Agents — the semantic data ontology that grounds natural-language answers in what the business actually means, rather than letting the model infer meaning (and risk fabricating it) (Source: sources/2026-09-15-databricks-data-ontology-defined-the-context-layer-your-ai-agents-are-missing).

Stub page. First wiki ingest naming Genie Ontology as a distinct, named component (vs the broader Genie product and Genie Code pipeline-generation surface).

What it does

Genie Ontology supplies the agent with the right, authoritative context at the point of reasoning — which definition of a metric applies, where the trusted data lives, which relationships and calculations matter — so the agent spends less time exploring schemas, reading documents, trying queries, and reconsidering assumptions. It is the concrete implementation of the "model the head, learn the tail" principle: humans explicitly model critical KPIs and business terms, while an inferred layer learns additional definitions, rules, relationships, and authoritative sources from existing work (dashboards, queries, notebooks, certified data, documentation, usage).

Governance is built in

  • Permissions enforced during retrieval. The ontology "should never become a back door around existing permissions." Genie Ontology enforces access at retrieval time using the governance of the underlying sources, including Unity Catalog — so two employees can ask the same question and appropriately receive different answers based on what each is authorized to see (concepts/governed-agent-data-access).
  • Authority ranking. Certification, authoritative definitions, lineage, usage, expertise, and source provenance are used to distinguish the official revenue definition from a one-off calculation from six months ago.

Disclosed benchmark

On an internal Databricks benchmark of 28 real-world enterprise data-analysis questions:

Agent First-attempt correctness Relative speed
Genie with Ontology 84.5% ~2× faster
Strongest general-purpose coding agent (unnamed) 52.4% baseline

Stated explicitly as an internal benchmark, not a universal accuracy guarantee. It is offered as evidence for the principle that better enterprise context can matter as much as, or more than, more reasoning time.

Relationship to other Genie / Databricks components

  • Genie One / Genie Agents — the analytics agents Genie Ontology grounds. The prior mechanism-level Genie disclosure names a "rich semantic enterprise context" derived from workspace assets as Genie's substrate; Genie Ontology is that context layer, named and governed.
  • Metric Views — governed metric/semantic layer that supplies authoritative measure definitions (the "head" humans model explicitly).
  • Unity Catalog — the governance perimeter whose ACLs the ontology honors at retrieval.

What's not disclosed

  • The inference mechanism for "learning the tail" — what signals feed it, how definitions/relationships are inferred, how authority is scored, how staleness is bounded.
  • Storage substrate (graph DB? flattened index? RDF?) — the post is definitional, not an internals disclosure.
  • Benchmark composition and the unnamed coding-agent baseline.
  • Latency / cost per query for the ontology-grounded path.

Seen in

Last updated · 766 distilled / 2,225 read