Skip to content

SYSTEM Cited by 3 sources

Unity Catalog Volumes

Unity Catalog Volumes are governed, versioned object-storage locations registered as first-class catalog assets inside Unity Catalog. They are the place non-tabular files (PDFs, images, audio, video, ML model artifacts) live when you want them inside the same governance surface as Delta tables.

Stub page. One ingested source so far.

Architectural role

Volumes are how the MapAid groundwater pipeline handles its raw input: the ~700 scanned PDFs / TIFFs / JPGs (>5,000 pages) plus the per-page rendered images all live in Unity Catalog Volumes. "Each document's pages are rendered as images and stored in Unity Catalog Volumes, creating a clean, versioned foundational dataset." (Source: sources/2026-05-11-databricks-unlocking-the-archives)

The pipeline then reads from those Volumes via ai_query — multimodal AI Functions can take Volume-stored image references as input columns. This composition (governed object storage → SQL-callable multimodal inference) is what lets the document-classification pipeline run as SQL/DataFrame jobs without standing up a separate file-handling service.

Why Volumes (not raw cloud-object-store)

  • Governance parity with Delta tables. Permissions, lineage, audit on the raw files live in the same UC surface as the output Delta tables.
  • Versioning. "Clean, versioned foundational dataset" — the raw archive is treated as a versioned snapshot, not a mutable bucket.
  • Pipeline-portability. The Asset Bundle references a Volume by name; pointing the bundle at a different archive is a config change, not a code change.

Seen in

  • sources/2026-08-28-databricks-fast-fault-tolerant-pytorch-training-on-ai-runtime — Governed store for both training checkpoints and training data. On AI Runtime, UCVolumeWriter/UCVolumeReader implement PyTorch Distributed Checkpoint against UC volumes — staging checkpoint I/O through local NVMe and marking a save complete only once data lands. Separately, UCVolumeDataset streams training files from a UC volume surfaced as a network mount, caching each to local NVMe on first access (see local-nvme-cache-for-remote-training-data). Two performance facts about the network-mount shape surface here: reading directly on every access binds step time to network latency, and it re-downloads the same files every epoch — which is why the local-NVMe cache in front of the Volume matters. The Volume keeps both checkpoints and training data inside the governed UC surface while local NVMe supplies the read/write speed.

  • sources/2026-08-10-databricks-how-to-ground-genie-agents-in-both-structured-data-and-documents-without-losing-governance — Governed knowledge source for Genie Agents. Landing documents in a Volume (rather than isolated storage with separate ACLs) puts unstructured knowledge inside the same governance plane as Delta tables; GRANT READ VOLUME then governs it under the same end-user-credential contract Genie uses for tables. Two Volume-specific facts surface here: (1) an attached Volume is a required source — a user lacking READ VOLUME on it "can't use that agent at all," so the grant gates the agent, not just document visibility; (2) a Volume is the smallest securable unit — all-or-nothing per Volume, no per-file scoping — which forces a one-audience-per-volume / one-agent-per-audience layout. Supported formats: PDF, JPG/JPEG/PNG/ TIFF/TIF, DOC/DOCX/PPT/PPTX, plain text, Markdown. See volume-as-agent-knowledge-source-with-required-access.

  • sources/2026-05-22-databricks-how-world-bank-group-uses-databricks-to-eradicate-poverty-through-shared-knowledge — RAG-corpus substrate for a multi-domain knowledge platform. World Bank Group indexes "tens of millions of documents" (project documents, publications) into UC Volumes paired with Vector Search to power a retrieval-augmented-generation capability — "using Databricks Volumes and vector search, they indexed project documents to create a retrieval-augmented generation capability that could respond to natural language queries and thus save manual search." Operational scale: 3M document downloads / month through the AI-powered search-and-synthesis layer, half from low- and middle-income countries. Caveat: chunking strategy, embedding model, indexing cadence, and per-document metadata schema not disclosed. The Volumes-as-RAG-corpus shape generalises the prior MapAid pattern (Volumes-as-multimodal-input-substrate) to a Volumes-as-document-corpus-fronting-Vector-Search composition.

  • sources/2026-05-11-databricks-unlocking-the-archives — canonical wiki instance. Stores raw scanned PDFs + per-page rendered images as the input substrate for the multimodal classification pipeline.

Last updated · 766 distilled / 2,225 read