Skip to content

SYSTEM Cited by 1 source

AI Extract Precision Mode

AI Extract Precision Mode is a high-accuracy mode of Databricks' ai_extract document-extraction API (a Databricks AI Function), aimed at the hardest enterprise extraction workloads: long documents with cross-page references, large deeply-nested outputs, and reasoning-heavy schemas. It is invoked by setting mode to precision in the ai_extract SQL function, or via the precision toggle in the Information Extraction UI on the Agents page. (Source: sources/2026-08-18-databricks-databricks-document-intelligence-pushing-the-frontier-for-complex-document-extraction)

Stub-plus page: documented from a single launch-post disclosure. Harness internals and the custom-model training recipe are not disclosed.

What it targets

Three extraction problems where a single frontier-model call breaks down:

  • Long documents — a lease whose page-1 renewal terms depend on a page-80 clause; docs up to 2,000 pages. Existing solutions fail to resolve cross-references.
  • Large, nested outputs — a multi-page bill of lading with hundreds of SKUs; invoices with thousands of line items; schemas with 300+ deeply nested fields. Existing solutions drop or truncate fields as outputs grow.
  • Complex schemas and reasoning — a risk classification synthesizing three financial statements; a contract value applying listed discounts across every recorded price.

Architecture (as disclosed)

Two layers combine to reach accuracy:

  1. Custom-trained extraction models. Rather than relying on ever-larger general-purpose models, Databricks trained models specialized to find, reason over, and extract structured information from complex documents, optimized for accurate structured extraction rather than open-ended generation.
  2. An agent harness — "inspired by Databricks MemEx" — that semantically decomposes large extraction jobs, executes smaller tasks in parallel, preserves intermediate results, and reconciles them into one final structured output. This is the semantic decompose → parallelize → preserve → reconcile pattern.

The harness is explicitly designed to survive model failure modes that break the naive baselines — the same reasoning as harness holds structure, agent holds reasoning: deterministic infrastructure enforces decomposition, intermediate-result persistence, and schema-conforming reconciliation; the model does the per-sub-task extraction and reasoning.

Why it beats chunk-and-merge

The realistic baseline is chunk-and-merge — split, extract per chunk, merge — not a single call. On long-document workloads frontier chunk-and-merge exhibited "chunk timeouts, truncated outputs, and incomplete final merges that did not conform to the requested schema." Precision Mode's differences:

  • Semantic decomposition (structure-aware, closer to smart chunking) vs. flat token-window splitting.
  • Preserved intermediate results enabling a reconciliation step that resolves cross-page references independent-chunk extraction cannot.
  • Reconciliation as a first-class stage producing one schema-conforming object, vs. an afterthought "merge" that fails schema conformance.

Reported performance

  • 94.7% accuracy across ~9,000 complex documents; +7 points over the strongest frontier chunk-and-merge baseline (GPT-5.6 Sol).
  • Benchmarks: 10 internal datasets (financial services, manufacturing, healthcare)
  • 5 public (VAREX, RealDocBench, LongExtractBench, LEDGER, Caselaw Access Project long-doc stress test).

(Vendor-reported; no independent reproduction — see source caveats.)

  • systems/databricks-ai-functions — ai_extract is one of these SQL-native primitives
  • systems/meta-harness — MemEx, the cited inspiration
  • systems/databricks
  • semantic-decompose-parallelize-reconcile-extraction — the harness shape
  • chunk-and-merge-extraction-baseline — the baseline it outperforms
  • smart-chunking — semantic-decomposition sibling
  • multi-step-llm-extraction — multi-stage LLM extraction concept

Seen in

Last updated · 766 distilled / 2,225 read