Databricks¶
Databricks Engineering blog. Tier-3 source on the sysdesign-wiki: most posts are product/marketing/ML-methodology oriented and get skipped, but infra-architecture posts (Kubernetes, service mesh, data-platform internals) are worth ingesting when they appear.
Internally Databricks runs hundreds of stateless gRPC services per Kubernetes cluster across thousands of clusters in multiple regions, predominantly in Scala on a monorepo with fast CI/CD. That monoculture is the architectural enabler for several of their infra-platform choices — notably the proxyless service-mesh design.
Key systems¶
-
systems/lakebase-vector / systems/lakebase-text — Lakebase Search (GA 2026-09-28): Postgres extensions for ANN vector search and BM25 full-text on Lakebase. lakebase_vector uses hierarchical IVF + binary RaBitQ quantization over object storage (a pgvector/ HNSW replacement that scales to zero and doesn't require the index in RAM); together they give native hybrid search inside Postgres. Benchmarked on VectorDBBench 100M.
-
systems/concurrence — Concurrence (first wiki disclosure 2026-09-23): a customer clinical-AI platform running agentic systems at ~1.2T annualized input tokens on Databricks. Notable for an event-sourced clinical "world model" (concepts/log-as-truth-database-as-cache) derived by Spark Declarative Pipelines from events ingested via Zerobus, served by Lakebase, governed per-tenant by Unity Catalog; compliance-gated model routing through Unity Gateway (BAA-namespace endpoint resolver); coding agents on the
ugCLI. A customer proper noun documented for its architecture, not a Databricks product. -
systems/fastmcp — FastMCP (first wiki disclosure 2026-09-25): a Python framework for composing MCP servers (
FastMCP.as_proxy(...)+mount(prefix=…)). At S&P Global Energy it is the engineering-owned composition layer that mounts per-dataset-group Genie Agent managed MCP servers behind per-commodity composite endpoints with name-spaced tools — the MCP-as-proxy pattern at the analytics altitude. Third-party framework documented for its role in a Databricks-customer architecture, not a Databricks product. -
systems/databricks-genie-one — Genie One + Genie One MCP (first wiki disclosure of the MCP server 2026-09-22). Genie One is the data-smart AI coworker for business users, powered by Genie Ontology; Genie One MCP exposes its governed conversational analytics as a five-tool MCP server (
genie_ask/genie_poll_response/genie_get_query_result/genie_cancel_response/view_ask) to any MCP-compatible agent (ChatGPT, Claude, Copilot, Claude Code), with OBO carrying Unity Catalog row/column governance into the user's context, MCP Apps interactive views, and a 90-second SQL timeout + workspace QPM limit. - systems/databricks-genie-ontology — Genie Ontology (first wiki disclosure 2026-09-15). The automatic context layer (data ontology) under Genie One / Genie Agents: humans model the critical-concept "head," an inferred layer learns the "tail" from existing dashboards/queries/notebooks/usage, permissions are enforced at retrieval via Unity Catalog, and knowledge is ranked by authority (certification / lineage / usage). Internal benchmark: 84.5% first-attempt vs 52.4% for the strongest coding agent (28 questions), ~2× faster.
- systems/databricks-radar — RADAR (Reliability Anomaly Detection, Alerting, and Root-cause analysis; first wiki disclosure 2026-09-19). Internal reliability-monitoring system that catches gray failures via streaming SPOT anomaly detection (Extreme Value Theory, 14-day window, single risk parameter) over per-error / per-region distinct-user-count series; enrich→filter→dedupe→routed ticket + Genie-backed RCA dashboard; ships as one DAB. 95% cut in incident-discovery time at >90% precision (vendor-reported).
- systems/proteus — Proteus (first wiki disclosure 2026-09-04). Agentic GPU-kernel-generation harness that produces inference kernels specialized to runtime operation shapes (1.8–5.2× over vLLM on Qwen 3.5 122B, Triton / B200). Its architecture centers on trusted validation and a scoped knowledge layer rather than the generator.
- systems/neonvm — NeonVM (first wiki disclosure 2026-08-31). QEMU/KVM Kubernetes custom resource + controller (Neon-origin) that resizes a running Postgres VM in place and live-migrates it (keeping its IP) — the mechanism behind Lakebase compute autoscaling. Lakebase's autoscaler tracks three signals (CPU, memory, and compute-cache working set — largest wins), estimates the current working set with a time-windowed HyperLogLog, and scales down as aggressively as up (symmetric-up-down-autoscaling). A production database can change size >32,000×/month.
- systems/lakehousert — LakehouseRT (first wiki disclosure 2026-08-27). Real-time low-latency analytics engine (codename Reyden) that queries open lake storage directly; the analytical half of the "first true" LTAP stack alongside Lakebase.
- systems/autoliquid — AutoLiquid (first wiki disclosure 2026-08-27).
Autonomic data-layout optimizer behind
CLUSTER BY AUTO: heuristic clustering-key selection + shadow verification across millions of tables, beating human-chosen keys on >95% of evaluated workloads (VLDB 2026). - systems/ultron — Ultron (first wiki disclosure 2026-08-27). History-based query optimization framework exploiting the repetitiveness of analytical workloads to improve optimizer choices (e.g. join operator); 25% median join-latency improvement (VLDB 2026).
- systems/pytorch-distributed-checkpoint — PyTorch Distributed Checkpoint (DCP) on
AI Runtime (first wiki disclosure 2026-08-28).
UCVolumeWriter/UCVolumeReaderimplement DCP against UC Volumes, staging through local NVMe withasync_save; the fault-tolerance backbone for high-goodput PyTorch training. Measured 58× faster checkpointing thantorch.savefor a 20B FSDP model on 32×H100. -
systems/databricks-ai-sre — AI SRE (first wiki disclosure 2026-08-24). Databricks' AI-powered incident-investigation agent for a fleet of 100s of microservices on 1500+ Kubernetes clusters across 70+ regions and three clouds. Automatic triage launches three parallel investigation tracks the instant an incident fires — platform health checks, service-level analysis, and team-owned agentic runbooks — and correlates them into a grounded diagnostic summary before the engineer opens their laptop, plus an interactive natural-language debugging mode. Built as a four-tier layered platform (Primitives → API Layer → Core Engine → Application Layer). Trustworthiness from three principles: structured checks before LLM reasoning, traceable evidence, and graceful degradation; plus API-layer guardrails for bursty agent access. 150+ teams, 250+ WAU, 2,000+ investigations/day. Design thesis: context assembly (not root-cause insight) is 60–80% of investigation time.
-
systems/ai-extract-precision-mode — AI Extract Precision Mode (first wiki disclosure 2026-08-18). High-accuracy mode of the
ai_extractAI Function for the hardest extraction workloads (long docs up to 2,000 pages, 300+ nested-field schemas, reasoning-heavy synthesis). Combines custom-trained extraction models with an agent harness inspired by MemEx that semantically decomposes → parallelizes → preserves intermediate → reconciles (semantic-decompose-parallelize-reconcile-extraction), engineered to survive the failure modes (chunk timeouts, truncated outputs, non-conforming merges) of the chunk-and-merge baseline. Reported 94.7% accuracy (~9,000 docs), +7 pts over the strongest frontier chunk-and-merge baseline. Launch-post disclosure; harness internals not detailed. -
systems/metals-v2 — Metals v2 (first wiki disclosure 2026-08-11). Databricks' fork-and-rework of the open-source Metals Scala language server, adding first-class Java support and re-architected for a 26M-line Bazel monorepo (2.9M symbols, 142k+ files). Optimizes TTII (p50 8.7s) via three layers: a build-free mbt content-addressed index, compiler-backed pipelines (one Scala presentation compiler holding 24M lines; Java on
javac+ Turbine), and a metadata-first BSP querying 285k Bazel targets off the editor hot path. Apache 2.0; ships in Cursor/VS Code/Neovim; Stripe early adopter. Confirms the "most code written by agents" posture that also motivates Omnigent and the Unity AI Gateway. -
systems/databricks-network-config-delivery — Network Configuration Delivery (first wiki disclosure 2026-08-12). Supplies every serverless VM its network config (allowed destinations, private-link endpoints, Unity Catalog grants, Delta Sharing) at tens of millions of launches/day → billions of requests/day. Re-architected from synchronous on-critical-path aggregation (~5,000 ms p99, 99.8%) to an event-driven background pipeline that materializes a per-workspace snapshot served by a single read (snapshot pre-computation + reconciler backstop): 125 ms p99, 99.99%, −86% upstream calls, with static config fallback on upstream outage.
-
systems/databricks-fevm — Field Engineering Vending Machine (FEVM) (first wiki disclosure 2026-07-25). Internal self-service infrastructure provisioning platform built as a Databricks App (React + Python) with Terraform for multi-cloud resource provisioning (AWS/Azure/GCP), Lakebase for state tracking, and an MCP interface for agent-first workflows. Core pattern: use-case-based provisioning — engineers describe intent, system maps to configured environment. TTL-based lifecycle with Slack transparency. Scale: 5,000+ active users, 2,600+ active deployments across 3 clouds, ~1,200 requests/day burst (BuildCon).
-
systems/databricks-custom-model-serving — Custom Model Serving (first wiki disclosure 2026-06-11). Fully managed real-time inference platform for any MLflow model — from 2 MB scikit-learn classifiers to 70B LLMs — without customer-facing tuning knobs. Three structural properties: isolated K8s deployments per endpoint, automatic runtime selection (Gunicorn/vLLM/Triton), and the AutoPilot Pod Autoscaler. Operating envelope: 300K+ QPS, 99.99% availability, 10→10K QPS in <60s. Eliminates the ML Stack Tax.
-
systems/databricks-autopilot-pod-autoscaler — AutoPilot Pod Autoscaler (APA) (first wiki disclosure 2026-06-11). Custom Kubernetes controller implementing two-axis autoscaling: horizontal (request-based, 5s interval) + vertical (model-aware concurrency tuning, 30s interval). The two axes are coupled — vertical output feeds horizontal formula. Asymmetric in both axes: aggressive up, conservative down. Heart of Custom Model Serving.
-
systems/lakebase-scm-extension — Lakebase SCM Extension (first wiki disclosure as a dedicated page 2026-05-30). The open-source VS Code / Cursor IDE extension maintained by
databricks-solutionsthat synchronises a developer's git branch with a matching Lakebase database branch — automating the per-developer paired-branch pattern inside the IDE — and surfaces the Branch Diff Summary view that the 2026-05-29 evolutionary-database-development post cites as the canonical schema-diff format. The IDE-substrate-glue primitive that closes the loop on Jen's per-developer branching workflow; alternative to thedatabricks postgres create-branchCLI flow. Public GitHub repo, behaviour-only disclosure in the source so far; architectural depth deferred to the forthcoming Companion: Plugin Walkthrough post that the 2026-05-29 article forward-references. -
systems/enzyme-ivm — Enzyme (first wiki disclosure as a dedicated page 2026-05-30). The incremental-view-maintenance engine that powers SDP's
@dp.materialized_viewdecorator. Subject of the SIGMOD 2026 honorable-mention paper "Enzyme: Incremental View Maintenance for Data Engineering" (arXiv:2603.27775, presented by Ritwik Yadav at SIGMOD 2026 in Bangalore). Establishes the two-track incremental-processing architecture inside SDP: Enzyme on the materialized-view track, Structured Streaming on the explicit-streaming track, mix-and-match in one pipeline. Four novel claims over prior industrial IVM: (a) full MV-grammar coverage including joins + windows + aggregations + combinations; (b) non-deterministic function support (current_date(), AI functions); (c) multi-language MVs (Python + SQL); (d) cost-model-driven incrementalisation strategy (partition-level vs row-level updates per run, selective intermediate-result caching, plan-info -
prior-execution-stats inputs). Articulated thesis: MV-as-ETL primitive — "if MVs can be efficiently and incrementally maintained, it will significantly simplify ETL workloads which otherwise require writing complex custom code".
-
systems/iceberg-v3 — Apache Iceberg v3 (first wiki disclosure as a dedicated page 2026-05-29). Spec-level milestone for Apache Iceberg, GA on Databricks 2026-05-28: three new format-level primitives — deletion vectors (file-level row-delete representation accelerating updates / merges / deletes without rewriting data files), row tracking (stable per-row identity for efficient incremental processing), VARIANT type (standard semi-structured-data type) — applied across managed Iceberg, foreign Iceberg, and UniForm-enabled managed tables. Delta-side cross-format compatibility called out verbatim as architectural precondition for the forward-looking Iceberg-v4 + Delta-5.0 adaptive-metadata-tree alignment (format-co-evolution-iceberg-delta).
-
systems/iceberg-rest-catalog-scan-api — Iceberg REST Catalog Scan Planning API (first wiki disclosure as a dedicated page 2026-05-29). The Iceberg-1.11 client API surface that Unity Catalog uses to extend ABAC across the engine boundary. Server-side scan planning — the catalog evaluates ABAC policies during plan-scan and returns a filtered scan plan; engines read only authorised data. Compatible engines: any implementing the Iceberg-1.11 scan-planning client (Spark, DuckDB named in the announcement). Wiki's first canonical instance of
scan-planning-as-policy-enforcement-point.
-
systems/databricks-axon — Axon (first wiki disclosure as a dedicated page 2026-05-29). The Databricks LLM data-plane router, named publicly for the first time in the 2026-05-27 Reliable LLM Inference at Scale post: "the data plane runs a router, which we call Axon, that balances load among replicas of the same model." Built on Dicer. Two structural properties distinguish it from the EDS+P2C path documented for the rest of Databricks Model Serving: load metric is model units (not active request count), and stateful (sticky) sessions route a workload's requests to a Dicer-assigned subset of pods — serving prefix-cache locality and blast-radius bounding simultaneously. Sits between rate-limiting and the inference runtime in the data plane. Production scale: 125T+ tokens/month across frontier OS (Kimi, Qwen) + proprietary (OpenAI, Gemini, Claude) models. The wiki's first canonical instance of patterns/ai-gateway-provider-abstraction + stateful-llm-session-routing.
-
systems/neon — Neon (first wiki disclosure as a dedicated page 2026-05-29; prior recurring tag mention across systems/lakebase / systems/pageserver-safekeeper). Serverless, separated-compute-and-storage Postgres, acquired by Databricks in 2025; architecturally identical to systems/lakebase (the Databricks-branded packaging). Operates the neonstatus.com public status surface. The 2026-05-27 reliability roadmap discloses Neon's empirical session-lifetime distribution: 90% of compute sessions for auto-suspending databases are <10 min — the load-bearing signal for the control plane is the new data plane reframe. Engineering lineage: Stas Kelvich (Neon co-founder) → Postgres internals expertise (multi-master replication with quorum commit, cross-node snapshot isolation under loosely-synchronised clocks).
-
systems/sqlsmith — SQLsmith (first wiki disclosure 2026-05-29). Open-source random SQL query generator paired with SQLancer in systems/lakebase's release-gate validation harness. SQLsmith finds crashes / assertion-violations under random query loads; SQLancer finds logic bugs against an SQL-standard oracle. Both run while fault injection is running so postcondition violations are detected during chaos drills. Verbatim from the source: "We utilize open source tools like SqlLancer and SqlSmith, along with similar internal tools, to verify correct Postgres behavior." See continuous-fault-injection-in-production.
-
systems/databricks-metric-views — Metric Views (first wiki disclosure 2026-05-27). UC-resident headless-BI semantic layer: define metrics ONCE in Unity Catalog (SQL or UC Explorer UI), every consumer (AI/BI Dashboards, Genie, SQL notebooks, third-party BI tools) resolves the same
MEASURE()definition. Semantic metadata fields (display_name/comment/synonyms) double as AI grounding context for Genie — natural-language questions get mapped to the right measure / dimension via the metadata, "no custom prompts, no separate glossary". Materialization is a substrate property (auto pre-aggregation + incremental refresh + intelligent query rewriting + transparent routing) — collapses three coupled artifacts (aggregate tables + refresh pipelines + BI-tool query updates) into one governed primitive. Open-standard provenance: SPARK-54119 (Apache Spark Metric Views OSS implementation) + UC OSS support coming. Canonical instance of headless-bi-semantic-layer + metric-view-materialization + governed-metric-as-headless-bi-substrate + auto-materialized-aggregation-via-semantic-layer + query-rewrite-to-pre-aggregated-materialization. - systems/databricks-predictive-optimization — Predictive
Optimization (first wiki disclosure as a dedicated page
2026-05-27; prior recurring tag mention only). Default-on
for UC managed tables;
automatically runs
OPTIMIZE,VACUUM, and statistics collection on tables that would benefit, "so you don't need to schedule these jobs yourself". Inline-during-Photon-writes collection of two statistics planes — Delta data-skipping statistics (per-file min/max/null for file pruning) + query optimizer statistics (cardinalities + distributions for plan choice). Back-fills stats for existing tables. 22% average performance improvement in observed workloads; "For BI workloads with repetitive filter patterns, the impact is especially significant". Drives theCLUSTER BY AUTOoption on Liquid Clustering — workload-aware automatic cluster-key selection. Canonical instance of automatic-table-optimization - optimizer-statistics-as-skipping-substrate.
- systems/databricks-sql-warehouses — DBSQL Warehouses
(first wiki disclosure as a dedicated system 2026-05-27).
Compute layer for BI / analytical SQL queries on the
Lakehouse; serverless auto-scaling (absorbs concurrency
bursts, pay-per-use); two-tier cache hierarchy —
disk cache (warehouse-local SSD for hot Parquet files) +
Query Result Cache (QRC) (full results keyed on query
text + table version, served without re-execution). For BI's
repetitive query patterns: "caching turns many requests into
millisecond-latency responses at near-zero compute cost".
Reflexive observability surface:
system.billing.usage+system.query.historysystem tables. Canonical instance of dbsql-caching-tiers. - systems/databricks-ai-bi-dashboards — AI/BI Dashboards
(first wiki disclosure 2026-05-27). First-party Databricks
dashboard surface; one of four named consumers of
Metric Views alongside
Genie / SQL notebooks / third-party BI tools. Source's worked
example: warehouse-metrics dashboard queried from
metv_dbsql_metricsMetric View (reflexively monitoring the warehouse via the same primitive the warehouse serves). - systems/octopus-margin-data-pipeline — Octopus Margin Data Pipeline (first wiki disclosure 2026-05-23). Customer-built Databricks-resident three-stream grain-aligned data pipeline for Octopus Energy's margin / settlement / commercial-KPI calculations under the UK MHHS regulatory transition (2 reads/customer/month → 48 reads/day, 48× volume). Settlement (HH) + Half-Hourly (smart tariffs: EVs, heat pumps, ToU) + Monthly (standard tariffs) on a unified multi-terabyte HH-grain source-of-truth substrate, orchestrated by "Job of Jobs". Single highest-leverage move: Delta CDF on the substrate — 25 B → 300 M rows/run (98.8%), weekly → daily freshness. Operating envelope: $0.48/settlement date (~50× below the projected MHHS cost, 2× below the legacy despite 48× more data), ~$1M/yr cost avoidance (excludes upstream-incremental savings — full gain is larger), 3 months, team of three. Canonical instance of grain-misalignment + grain-aligned-stream-split + job-of-jobs-orchestration + broadcast-join-for-small-reference-tables (<500 MB threshold disclosed) + remove-before-add-optimization (Z-ordering / ANALYZE / custom shuffle named as audit targets; AQE outperformed hand-tuning and the team deleted code).
- systems/spark-aqe — Spark Adaptive Query Execution (first wiki disclosure as dedicated page 2026-05-23; prior recurring tag mention only). Runtime query optimiser inside Spark — re-plans during execution using post-shuffle row counts, observed skew, and post-filter join input sizes (statistics the static planner doesn't have). The Octopus rebuild canonicalises the trust-the-optimiser-as-an- architectural-move face: "In several cases, Spark's Adaptive Query Execution (AQE) outperformed hand-tuned logic. The team removed custom optimisation code and let AQE do its job." The action item is deletion, not addition. Pairs with remove-before-add-optimization as the generalised principle.
- systems/liquid-clustering — Liquid Clustering (first wiki disclosure as dedicated page 2026-05-23; prior recurring tag mention across systems/uc-otel-trace-tables, systems/uc-managed-tables, systems/zerobus-ingest, systems/delta-lake, concepts/observability, patterns/telemetry-to-lakehouse, patterns/telemetry-to-lakehouse). Delta Lake feature that "dynamically co-locates related records on the specified clustering keys without requiring fixed partition boundaries" — the partition-replacement primitive. The Octopus disclosure enumerates the over-partitioning failure modes verbatim: "Liquid clustering avoids the small-file problem, higher memory consumption, and I/O overhead that come from over-partitioning." Pairs with broadcast-join-for-small-reference-tables in the "join and partition tuning" category.
- systems/zerobus-ingest — Zerobus Ingest (first wiki disclosure 2026-05-22; deep architecture disclosure 2026-06-11). Managed serverless push-based streaming ingestion engine. The 2026-06-11 post reveals three key internals: (a) dynamic partitioning with stream-connection-level ordering (not partition-level) enabling true elastic autoscaling; (b) Zeroparser — Rust-based zero-copy protobuf decoder achieving ~1 GB/s/core with dynamic descriptors, outperforming codegen (OSS); (c) latency-optimized WAL with async offset acks via gRPC bidirectional streaming, Delta Kernel Rust for final Delta commit. Benchmark: 12 GB/s sustained / 12M rows/sec / 1.04T rows in 24h from 2,048 concurrent streams. Also serves as the OTel-protocol ingestion endpoint (OTLP/gRPC + REST) for the single-sink telemetry architecture; canonical instance of stream-connection-as-ordering-unit + zero-copy-protobuf-decoding + wal-before-lakehouse-publish.
- systems/uc-otel-trace-tables — UC OTel Trace Tables
(first wiki disclosure 2026-05-22). Six MLflow-derived UC
Delta views —
<prefix>_otel_spans(per-request span execution),_otel_logs(structured log/event data),_otel_metrics(numerical telemetry),_otel_annotations(MLflow-specific tags / assessments / feedback / expectations / run-links),_trace_unified(one record per trace with raw spans + metadata),_trace_metadata(MLflow trace metadata grouped by trace ID — "more performant than the unified view when you only need MLflow trace metadata"). Auto liquid- clustered post the latest product update; MLflow per-experiment trace cap eliminated. Inherits UC governance for free — column masking on prompt/response columns, row-level filtering by tenant tag, RBAC, audit logs. Sibling lakehouse audit substrate to Inference Tables at a different granularity (one row per span vs one row per model call). - systems/mlflow-otel-tracing — MLflow OTel Tracing
(first wiki disclosure 2026-05-22). Framework-side OTel
instrumentation surface within MLflow 3 —
mlflow.<lib>. autolog()for automatic per-call span emission (LangChain, LangGraph, OpenAI SDK, etc.) +@MLflow.tracedecorator for the request-level root span. Provisions the UC OTel trace tables from Python on experiment setup; same code path whether the agent runs inside Databricks, in the customer's VPC, on a developer laptop, or in a third-party cloud ("In fact the support assistant agent example that was used for this blog is deployed locally"). Substrate for the prod-traces-as-eval-dataset flow (see production-traces-as-evaluation-substrate + bootstrap-eval-dataset-from-production-traces) — same judges run on bootstrapped historical-prod traces in dev and on live traces in production. - systems/databricks-fmapi-prompt-caching — FMAPI Prompt Caching (GA on open-weights models 2026-05-22). Implicit, volatile-only, tenant-isolated KV-cache reuse shipped as a default-on substrate property of the Foundation Model APIs; every layered product (Agent Bricks, Genie, AI Functions) inherits the win at no integration cost. Disclosed numerical signature on the GPT-OSS batch-inference rollout: +2.5× per-replica throughput, 3× P50 latency reduction at 30% cache hit ratio. Catalog: GPT-OSS 20B + 120B, Gemma 3 12B, Llama 3.1 8B (incl. PEFT-served fine-tuned variants), Llama 3.3 70B.
- systems/uc-managed-tables — Unity Catalog Managed Tables (first wiki disclosure 2026-05-14). Open-API-accessible managed Delta tables; Predictive Optimization + Liquid Clustering produce "up to 20× faster queries and 50% lower storage costs"; external engines (Apache Spark, Apache Flink, DuckDB) can now create, read, write, and stream to/from these tables via Delta Kernel; safety substrate is UC catalog commits (serialized commits + complete auditability + multi-table- transaction coordinator). Beta version pinning: Delta-Spark 4.2 + UC 0.4.1.
- systems/uc-credential-vending — Unity Catalog Credential Vending API (GA for tables 2026-05-14, Public Preview for Volumes). Mints short-lived, scoped credentials on demand for external engines; M2M OAuth replaces personal access tokens (named "per-user, long-lived, and hard to rotate"); engines auto-refresh via the vending API so "pipelines that run for hours complete reliably without tokens expiring mid-job." Same primitive extends to UC Volumes for unstructured assets (images / PDFs / videos). Canonical instance of concepts/short-lived-credential-auth.
- systems/delta-kernel — open-source Java + Rust library for reading, writing, and committing to Delta tables (first wiki disclosure 2026-05-14). Abstracts the Delta protocol so "connector developers can focus on UC integration, not Delta implementation." Three named adopters: Apache Spark (via Delta-Spark 4.2), Apache Flink (via Delta Flink), DuckDB. Canonical instance of connector-library-as-protocol-abstraction — the ecosystem-growth lever that lets any engine integrate with Unity Catalog at low cost.
- systems/databricks-apps — the workspace-resident application runtime (first wiki disclosure 2026-05-13). Web app deployed inside the Databricks workspace; authenticates as a workspace service principal; queries Unity Catalog via the SQL Statement API; calls AI/BI Genie over the workspace REST API — "all on internal connections". Composes with systems/lakebase (operational app state, scale-to-zero) + UC (governed analytical data + ML audit substrate) + Genie (embedded NL query) + systems/mlflow (model + attribution versioning) into the single-platform application architecture thesis. Reference implementation: systems/site-feasibility-workbench (FastAPI + React, ~30 min deployment; clinical-trial site selection under FDORA 2022 + 21 CFR Part 11 + ICH E6(R3) + FDA GMLP). Forward roadmap: three additional Databricks Apps (Patient Cohort and Recruitment, Enrollment Velocity Optimizer, Risk-Based Monitoring and Compliance) on the same shape.
- systems/site-feasibility-workbench — first public open-source Databricks App (released 2026-05-13). FastAPI backend + React frontend; six-step guided clinical-trial site-selection workflow; TA-segmented LightGBM models trained on sponsor CTMS / EDC / IRT history; per-recommendation SHAP attributions written to UC-governed Delta tables (substrate + pattern). Saved shortlists persist to systems/lakebase; AI/BI Genie embedded for cross-domain NL-query against the same UC tables the ML models trained on. Reference implementation for single-platform-application-architecture and in-workspace-app-as-decision-support.
- systems/databricks-ai-functions — the SQL-callable LLM
inference primitive (
ai_query) used as the universal inference surface in the 2026-05-11 MapAid groundwater pipeline ingest. Multimodal input (page images), structured JSON output, and three load-bearing roles in one pipeline (classification / extraction / judge) without a separate model-serving service. Canonical instance of SQL-native multimodal LLM inference. - systems/databricks-foundation-model-api — managed multimodal model endpoint behind AI Functions; serves the full-page OCR + entity-recognition pass on water-flagged documents in the MapAid pipeline.
- systems/databricks-asset-bundles — declarative pipeline-packaging unit. The MapAid groundwater pipeline ships as one bundle deployable + runnable with one command; pipeline-vs-archive decoupling makes it portable to other domains (other water archives, regions, scanned-document corpora). Canonical instance of asset-bundle-single-command-deployment.
- systems/lakeflow-jobs — orchestration layer running the MapAid multi-stage pipeline on serverless compute.
- systems/unity-catalog-volumes — versioned, governed object storage for non-tabular assets (raw scanned PDFs/TIFFs/JPGs and per-page rendered images in the MapAid pipeline). First wiki disclosure of UC's Volumes face.
- systems/databricks-model-serving — the managed real-time
inference platform disclosed at platform-internals depth in
the 2026-05-08 Databricks / Superhuman joint post. Two-layer
co-engineered stack: platform layer (EDS-driven
Power-of-Two-Choices LB +
request_concurrency-based asymmetric autoscaler + lazy-loading container image) and runtime layer (FP8 quantisation with per-channel scaling + hybrid- precision toggle + multiprocessing RPC runtime for the CPU-bound regime + async CPU-GPU scheduler). Operating envelope (Superhuman workload): 200,000+ QPS peak, sub-1-second p99, 4-9's reliability with per-pod throughput 750 → 1,200 QPS on H100 (+60%) for a ~50/50-token-shape LLM. Canonical instance of the managed- serving-without-giving-up-control split: customer owns model + quantisation + quality bar; platform owns runtime + LB + autoscaler + image substrate. - systems/databricks-endpoint-discovery-service — the xDS
control plane Databricks built for intra-cluster Armeria RPC
(2025-10-01) and promoted to managed external inference at
200K+ QPS via the 2026-05-08 Superhuman migration. Watches
Kubernetes API for
Services+EndpointSlices, streams endpoint state to clients implementing P2C. Same control plane, expanded blast radius across two altitudes (intra-cluster service mesh + external inference). - systems/superhuman-grammar-correction-model — Superhuman's custom LLM serving real-time grammar / clarity / tone / style suggestions across the Superhuman productivity suite (Coda, Mail, Go) at peaks of 200,000+ QPS for 40M+ daily users with sub-1s p99 and 4-9's reliability. Canonical wiki instance of a small-fast-LLM at massive QPS — and therefore the canonical instance of the CPU-bound serving regime on H100. Pre-migration stack: DIY vLLM on L40S; post-migration: Databricks Model Serving on H100.
- systems/claroty-cps-library — Claroty's AI-Powered CPS Library, the asset-identity layer behind the xDome cyber- physical-systems protection platform (2026-05-13 disclosure). Catalog scale: 17 million+ assets consolidated as canonical CPS-IDs from noisy plant-floor evidence (protocol-derived model strings, vendor codes, firmware markers, OEM PDFs) into vulnerability records (CVEs / CISA advisories / NVD / CPE). Hybrid Entity Resolution architecture composing Databricks primitives: Medallion over Delta Lake with Delta CDF driving a versioned mapping registry; orchestrated multi-agent system (NLP / Reasoning / HITL) on Model Serving custom endpoints + Knowledge Assistant + Information Extraction; MLflow continuous evaluation via LLM-as-a-Judge against concept drift; Lakebase as transactional asset-mapping store with strict constraints; Apps
- Lakebase as the SME human-in-the-loop UI;
Lakeflow Jobs orchestrating CSAF
security-advisory ETL with
ai_queryper-step. Reported outcomes: 25% improvement in vulnerability identification accuracy; 56% of analysed devices receive new or updated security recommendations for previously-invisible outdated firmware. Canonical instance of hybrid-classical-er-plus-genai + orchestrated-multi-agent-entity-resolution. - systems/databricks-serverless-compute — the serverless Apache Spark product framed under the thesis "stability becomes a system property rather than a user responsibility" (2026-05-06). Composes three systems: Spark Connect (gRPC driver-client split), Serverless Gateway (three-signal workload-aware routing), and Serverless Autoscaler (two-axis adaptive scaling with OOM-aware VM restart). Scale: 25+ major Spark runtime upgrades per year at 99.998% success across >2 billion workloads (per SIGMOD/ PODS '25 "Blink Twice" paper). Customer outcomes: CKDelta 20 min vs 4–5 hr, Unilever 2–5× faster + 25% cost reduction, HP 32% savings + 36% runtime reduction. Canonical instance of stability-as-system-property at Spark's compute altitude.
- systems/spark-connect — the gRPC client-server rearchitecture of Spark's driver. "The most significant architectural transformation in Spark's history" — user application code no longer co-executes with the driver; queries travel as serialised logical plans over gRPC. Unit of execution shifts from processes to queries. Canonical instance of grpc-decoupled-driver-client at the Spark altitude.
- systems/databricks-serverless-gateway — workload-aware router using three real-time signals (logical-plan-derived query size + cluster utilisation + interactive-vs-batch latency profile) with continuous re-evaluation. Resolves the utilization-vs-predictability-tradeoff at the pool layer. Canonical instance of patterns/ai-gateway-provider-abstraction.
- systems/databricks-serverless-autoscaler — two-axis adaptive autoscaler (horizontal + vertical). OOM-aware VM- restart primitive detects task-level OOM and re-executes the task on a larger VM without job failure. Canonical instance of adaptive-oom-recovery + concepts/elasticity + oom-aware-vm-restart-autoscaling.
- systems/pantheon — Databricks' internal fork of CNCF Thanos, the TSDB at the core of Databricks' global monitoring stack. 160+ instances across ~70 cloud regions on 3 major clouds; ~5 billion active in-memory timeseries and >10 trillion samples/day. Two Receive groups with distinct memory-retention tiers (2h for persistent services, 30m for ephemeral serverless) matched to workload lifespan. Purpose-built 3-controller control plane (Rollout Operator / Hashring Controller / Autoscaling + Self-Healing Controller) remediating dozens of incidents per week. At-least-once block uploads (2 of 3 StatefulSets upload) cuts object-storage egress. Migration saved "millions of dollars in annual cloud costs" with "~5× reduction in monitoring infrastructure downtime." Upstream contributions to Thanos.
- systems/hydra — Databricks' lakehouse-native observability platform for raw unaggregated high-cardinality troubleshooting data. Spark Structured Streaming + Databricks Auto Loader ingest 20 billion active unaggregated timeseries into Delta Lake with ~5 min freshness. ~50× cheaper storage than Thanos. PromQL-to-SQL translation layer lets Grafana + existing dashboards query Delta tables unmodified. Unified metric semantics across both paths (Pantheon TSDB + Hydra lakehouse) — engineers don't distinguish.
- systems/telegraf — InfluxData OSS agent, deployed as the cardinality shield in front of Pantheon. >1 GB/s per region, thousands of aggregation rules, built on Dicer sticky routing (not Kafka) to preserve in-memory aggregator state across redeploys. Absorbed a 2-5× metric surge during an infra incident so Pantheon only saw a 20% surge.
- systems/dicer — Databricks' open-sourced (2026-01) auto-sharder. Dynamic slice-range sharding with hot-key isolation + replication, eventually-consistent Assignments, state transfer across reshards. Used by Unity Catalog, Softstore, the SQL query orchestration engine, and "every major Databricks product". 2026-05-05 addition: powers the Telegraf metric-aggregation tier above via sticky-routing-for-aggregator-state.
- systems/unity-catalog — unified governance service; Dicer-backed sharded in-memory cache drove 90–95% hit rate and drastic DB-load reduction. Also the hub of customer-facing data meshes — federates external Iceberg catalogs and exchanges data via systems/delta-sharing. 2026-05-13 GA disclosure: also the policy-evaluation engine for the organize → detect → protect governance pipeline, hosting three co-designed primitives (governed tags + ABAC policies + agentic data classification) inside one permission + metadata model.
- systems/unity-catalog-abac — ABAC policy primitive
(GA 2026-05-13). Evaluates tag-based conditions to apply
row filters + column masks automatically across catalogs/
schemas. 10K+ policies per metastore, 100+ per catalog/schema
(10× growth at GA). Session identity evaluation for
views/functions closes the view-as-bypass failure mode. Single
VARIANT UDF can mask
INT/DOUBLE/DECIMAL/STRUCTcolumns at once. Customer testimonials: Atlassian (operational- overhead reduction), Udemy ("Fewer policies, lower costs, surgical precision"). - systems/unity-catalog-governed-tags — account-level governed tag taxonomy (GA 2026-05-13). Tags attach to catalogs, schemas, tables, columns; inherit parent → child; CREATE/MANAGE permissions distinct from data ownership; full SQL DDL + REST + UI
- Terraform lifecycle. The attribute foundation ABAC policies evaluate against and the output substrate Data Classification writes into.
- systems/unity-catalog-data-classification — agentic classifier (GA 2026-05-13; custom classifiers in Beta). Pattern-recognition + metadata + LLM signals continuously tag sensitive columns. Built-in classifiers cover GDPR / HIPAA / GLBA / DPDPA / PCI plus UK / Germany / Australia / Brazil regional packs (India + Canada coming May 2026). Custom classifiers learn from already-tagged columns. Human-in-the-loop FP exclusion improves precision over time. Output substrate is the same governed-tag vocabulary humans use.
- systems/delta-sharing — open cross-cloud / cross-metastore / cross-partner data-exchange protocol. Used by Mercedes-Benz for three deployment shapes (cross-hyperscaler, cross-region, external partner) on one wire protocol.
- systems/delta-lake — Databricks' open table format; Deep Clone is the incremental-replication primitive behind cross-cloud-replica-cache.
- systems/softstore — distributed KV cache built on Dicer; canonical example of Dicer's state-transfer (~85% hit rate preserved through rolling restarts vs. ~30% drop without).
- systems/databricks-endpoint-discovery-service — custom xDS control plane watching Kubernetes services/EndpointSlices, feeding both Armeria RPC clients (internal) and Envoy ingress (external) off one source of truth.
- systems/armeria — shared Scala RPC framework; host of embedded client-side LB + xDS subscription code.
- systems/storex — internal AI-agent platform for database debugging across the global fleet; central-first sharded architecture
- DsPy-inspired tool framework + snapshot-replay validation with judge LLMs.
- systems/dspy — Databricks-sponsored programmatic prompt framework; cited as inspiration for Storex's tool/prompt decoupling.
- systems/mlflow — Databricks-originated ML lifecycle platform;
hosts Storex's
judgesprimitive and prompt-optimization tooling. - systems/lakebase — Databricks' serverless Postgres (Neon lineage, 2025 acquisition); Pageserver + Safekeeper durable storage, ephemeral Postgres compute VMs. 2026-04-20 CMK rollout ingested; 2026-04-27 first production deployment (LangGuard) ingested.
- systems/langguard — runtime enforcement layer for agentic workflows, profiled 2026-04-27 as one of the first startups on Lakebase. Intercepts every agent tool/data/model invocation and returns allow/deny/modify synchronously. Built by the IBM QRadar (SIEM) team; canonical articulation of "database architecture is destiny" for bursty security- telemetry-shaped workloads.
- systems/grail-data-fabric — LangGuard's patent-pending governance engine; live knowledge graph of workflow behavior + context backing runtime policy evaluation. Runs on Lakebase.
- systems/pageserver-safekeeper — the Neon-lineage page + WAL durable storage tier Lakebase inherits.
- systems/aws-kms / systems/azure-key-vault / systems/google-cloud-kms — the three cloud KMSes Lakebase's Customer-Managed Keys feature integrates with.
- systems/unity-ai-gateway — productised AI-gateway for coding agents + MCP governance (launched 2026-04-17). Three pillars: centralised audit in Unity Catalog, single-bill cost control via Foundation Model API + BYO external capacity, OpenTelemetry → UC-Delta-table observability. Clients ready at launch: Cursor, Codex CLI, Gemini CLI, with Claude Code via MLflow 3 tracing.
- systems/databricks-foundation-model-api — first-party inference for OpenAI/Anthropic/Gemini/Qwen underneath Unity AI Gateway; BYO external capacity supported.
- systems/databricks-join-order-agent — UPenn-collaboration
research prototype (2026-04-22): frontier LLM agent as an
offline join-order tuner for the Databricks query engine.
Single tool (
execute_plan), 50-rollout budget, grammar- constrained structured output, best-of-N selection. On the JOB benchmark over 10× scaled IMDb: 1.288× geomean speedup, 41% P90 drop — outperforming perfect cardinality estimates, smaller LLMs, and the classical BayesQO baseline. - systems/join-order-benchmark-job — the canonical 113-query IMDb-based academic benchmark Databricks' join-order agent is evaluated on; scaled 10× via row-duplication for the Databricks experiment.
- systems/bayesqo — Bayesian-optimization-based offline query-plan optimizer (Postgres-oriented prior art) used as the comparison baseline.
- systems/databricks-glow — Databricks' open-source distributed genomics toolkit on Spark (VCF / BGEN / PLINK → Delta tables). Named 2026-04-22 as the genomics-modality ingestion tool inside the lakehouse-as-multimodal-substrate pattern.
- systems/lakeflow-spark-declarative-pipelines — Databricks'
declarative streaming-ETL layer;
@dp.table+@dp.materialized_viewdecorators onpyspark.pipelinesfor streaming tables with schema evolution + late events + continuous aggregation. Named 2026-04-22 as the wearables- streaming tool inside the same pattern. - systems/mosaic-ai-vector-search — Databricks' managed vector search over governed Delta tables; indexes imaging-derived feature embeddings for similarity queries ("find similar phenotypes within glioblastoma") without moving vectors out of Unity Catalog's governance domain.
- systems/databricks-autocdc — declarative
CDC /
SCD API inside
Lakeflow SDP:
dp.create_auto_cdc_flowwithkeys,sequence_by,apply_as_deletes,stored_as_scd_typeparameters. Replaces 40–200+ lines of hand-rolledMERGElogic with ~6–10 lines of declarative definition. Supports CDF sources, SCD Type 1 / Type 2, and snapshot-diff inference. Runtime gains since Nov 2025: 71% perf-per-dollar SCD Type 1, 96% SCD Type 2. Adopters include Navy Federal Credit Union (billions of events/day), Block, Valora Group. Named 2026-04-22. - systems/databricks-genie-code — Databricks' AI-assisted
pipeline-generation product. Positioned 2026-04-22 as the
LLM-codegen client that produces AutoCDC declarations rather
than raw
MERGElogic, so AI-generated pipelines inherit AutoCDC's bounded-correctness envelope. - systems/databricks-genie — Databricks' state-of-the-art data agent for natural-language analytics over enterprise data (structured + unstructured). 2026-04-29 disclosed empirically via Trinity Industries (>1,000 questions/month, three-stage adoption curve). 2026-05-08 disclosed at architecture-level: three named advances (specialised knowledge search with up-to-40% table-discovery benefit; parallel thinking via multi-trajectory sampling + aggregation as the structural response to the verifiable-test gap; Multi-LLM with per-sub-agent assignment + GEPA- optimised prompts). Headline: 32% → over 90% accuracy vs "a leading coding agent" on Databricks' internal benchmark, simultaneously on accuracy + cost + latency. Operates via the four-phase trajectory (discovery → investigation → self-correction → verification). Architecturally distinct from coding agents — see data-agent-unique-challenges.
- systems/gepa-prompt-optimizer — prompt-optimisation method (arXiv 2507.19457) referenced by Genie's Multi-LLM advance as the per-(LLM, sub-agent) prompt optimisation tool that closes accuracy gaps for smaller / faster sub-agent models. First wiki citation 2026-05-08.
- systems/apache-datasketches — Apache Software Foundation library of production-grade probabilistic data structures (KLL, Theta, approx top-K, Tuple). Databricks exposes it as first-class SQL / DataFrame / Structured Streaming aggregates in the 2026-04-29 launch, backing dashboards / analytics workloads that accept 1–2% relative error in exchange for orders-of-magnitude compute reduction. Community contribution: Christopher Boumalhab implemented the Theta + Tuple sketch function families in upstream Apache Spark.
- systems/databricks-postgres-cli — the
databricks postgres generate-database-credentialcommand family inside the broader Databricks CLI; mints scoped, short-lived OAuth JWTs for Lakebase endpoints. Canonical first wiki integration datum 2026-04-30: Thoughtworks Backstage POC used a 50-minute cron wrapping this command to bridge Backstage's long-lived-credential expectation against Lakebase's short-lived-JWT auth posture. - systems/backstage — Spotify's open-source Internal Developer Portal framework (CNCF incubating); its Postgres- default + Knex-migration + notoriously-fragile-schema shape makes it the representative state-heavy-application migration-stress-test for Lakebase. Thoughtworks POC 2026-04-30 migrated a Backstage install off standard Postgres onto Lakebase to demonstrate cheap branching + PITR at 63 MB catalog / 1.09-second branch / 3.78-second recovery.
-
systems/thoughtworks-technology-radar — Thoughtworks' twice-yearly industry-guidance publication; endorsed Backstage as IDP foundation, motivating the 2026-04-30 Backstage-on- Lakebase POC.
-
systems/gpu-monitor — gpu-monitor (first wiki disclosure 2026-07-05). Multi-stage health check and observability service running on every GPU node in the Databricks AI training fleet. Three layers: active bootstrap checks (GPU burn-in, peer connectivity, NCCL correctness, ECC/HBM, PCIe, DCGM L2), passive continuous checks (NVLink lanes, clock throttling, RDMA port down, XID errors, thermal gradients), and periodic multi-node probes (NCCL collective bandwidth sweeps 8B–2GiB). Implements multi-stage-health-check and node-quarantine-and-retest.
Key patterns / concepts¶
- shard-replication-for-hot-keys — isolate a hot key into its own slice, replicate that slice across N pods (Dicer's answer to concepts/hot-key).
- state-transfer-on-reshard — migrate per-key state across pods during assignment changes so caches survive rolling restarts (Softstore).
- dynamic-sharding — continuously-adjusted Assignment driven by health + load signals; Dicer's core primitive.
- concepts/horizontal-sharding / concepts/hot-key / no-distributed-consensus — the three structural failure modes Dicer was built to replace.
- concepts/eventual-consistency — Dicer's Assignment consistency model; trade-off vs. Slicer / Centrifuge leases.
- concepts/leader-election — key-affinity-as-coordinator, one of Dicer's named use cases.
- proxyless-service-mesh — mesh capabilities (discovery, L7 LB, health-aware routing, zone-affinity) via shared library instead of sidecars. Rejected Istio / Ambient Mesh explicitly.
- patterns/power-of-two-choices — the default LB algorithm embedded in Armeria clients.
- patterns/client-proximal-leader-pinning — with capacity/health-driven spillover.
- slow-start-ramp-up — introduced after client-side LB surfaced cold-start issues on fresh pods.
- patterns/specialized-agent-decomposition — DsPy-inspired; tools-as-functions with docstrings, prompts and tools swap independently.
- patterns/snapshot-replay-agent-evaluation — production-state snapshots replayed through candidate agent configs, scored by a judge LLM — Storex's regression harness.
- patterns/specialized-agent-decomposition — per-domain agents (DB, traffic, …) collaborating on root-cause analysis.
- hackathon-to-platform — 2-day prototype → user-feedback iterations → platform; Storex's on-ramp.
- concepts/client-side-load-balancing — overall architectural posture for internal RPC.
- layer-7-load-balancing — why they left kube-proxy.
- xds-protocol — the control-plane-to-data-plane contract, used beyond sidecar meshes here.
- concepts/control-plane-data-plane-separation — EDS vs. client/Envoy.
- concepts/tail-latency-at-scale — the observed symptom that motivated the proxyless redesign.
- central-first-sharded-architecture — Storex's foundation: global coordinator + regional shards for data-residency + one auth model across 3 clouds / hundreds of regions / 8 regulatory domains.
- concepts/llm-as-judge — scoring primitive inside Storex's validation framework.
- concepts/data-mesh — Unity Catalog + Delta Sharing position as the Databricks-opinion answer to the mesh shape, with domain-owned data products + central governance + open exchange protocol.
- hub-and-spoke-governance — the UC governance posture; central catalog + federated spoke data, one policy surface across clouds/regions.
- cross-cloud-architecture / concepts/egress-cost — the forcing functions behind the Mercedes-Benz mesh design.
- cross-cloud-replica-cache — Delta Sharing + Delta Deep Clone as the canonical shape for bulk cross-cloud consumers.
- chargeback-cost-attribution — bytes-at-the-sync-tier → producer-side billing dashboard; the governance hygiene layer on top of the replica-cache pattern.
- concepts/envelope-encryption — the three-level CMK → KEK → DEK hierarchy the Lakebase CMK rollout articulates cleanly.
- concepts/byok-bring-your-own-key — customer-held root of trust model; Lakebase realises it across both its storage and compute tiers.
- cryptographic-shredding — Lakebase revocation semantics: unwrap fails → data cryptographically inaccessible; Manager terminates compute VMs.
- per-boot-ephemeral-key — Postgres compute VMs generate a per-boot ephemeral key that dies with the instance; pairs with concepts/stateless-compute scale-to-zero.
- concepts/compute-storage-separation — Lakebase's Pageserver+ Safekeeper vs ephemeral Postgres compute split is a canonical OLTP-shape instance (cf. Aurora DSQL, Snowflake).
- coding-agent-sprawl — named problem class: engineering orgs simultaneously running Cursor + Codex + Claude Code + Gemini CLI + … ; Databricks itself is the stated example.
- concepts/centralized-ai-governance — three-pillar framing (security/audit + single bill + Lakehouse observability) that Unity AI Gateway instances, paralleling Cloudflare's internal-stack shape with different substrates.
- patterns/ai-gateway-provider-abstraction — Unity AI Gateway is the Databricks instance, specialised for coding-agent + MCP governance.
- patterns/central-proxy-choke-point — architectural posture: all coding-tool traffic funnels through one gateway, no second path to providers.
- patterns/unified-billing-across-providers — Foundation Model API as first-party default + BYO external capacity = one bill + per-developer (not per-tool) budgets.
- patterns/telemetry-to-lakehouse — coding-tool OpenTelemetry → UC-managed Delta tables; joinable with Workday / PR-velocity / capacity-planning data.
- join-order-optimization — the decades-old query-planner subproblem the 2026-04-22 research prototype targets.
- cardinality-estimation — named as "as difficult as executing the query itself"; the weakest leg of the three-component optimizer decomposition.
- llm-agent-as-query-optimizer — the architectural pattern the Databricks/UPenn agent canonicalises.
- offline-query-tuning-loop — the human-DBA- workflow shape LLM agents can automate.
- anytime-optimization-algorithm — the algorithmic family rollout-budgeted agents belong to.
- exploration-exploitation-tradeoff-in-agent-search — per-rollout allocation decision inside the agent.
- like-predicate-cardinality-estimation-failure —
canonical
LIKE-predicate blind spot illustrated by JOB query 5b. - llm-agent-offline-query-plan-tuner — full pattern: one tool, rollout budget, grammar-constrained output, best-of-N.
- structured-output-grammar-for-valid-plans — grammar constraints ensure every rollout lands on a semantically-legal plan.
- rollout-budget-anytime-plan-search — bound search by rollout count → monotonic best-so-far, time-budget knob.
- governed-delta-tables-per-modality — the named remedy to the specialty-store-per-modality anti-pattern: every modality (genomics, imaging, notes, wearables) lands in governed Delta tables under one Unity Catalog surface; modality-specific tooling layers above the substrate rather than defining a separate stack per modality.
- fusion-strategy-selection-by-deployment-reality — pick multimodal fusion (early / intermediate / late / attention) from deployment-reality axes (modality availability, dimensionality balance, temporal dynamics); late fusion is the "safe start" default when missingness is expected.
- missing-modality-problem — "missingness isn't an edge case — it's the default"; Databricks canonicalises sparse-modality deployments as a first-class sysdesign concern.
- modality-masking-during-training — training-time regulariser that drops modality inputs to simulate deployment sparsity.
- early-fusion / intermediate-fusion / late-fusion / attention-based-fusion — the four fusion strategies each paired with a deployment-reality trigger.
- declarative-cdc-over-hand-rolled-merge — declare
CDC / SCD semantics (keys, sequence column, delete predicate,
SCD type); runtime implements ordering, dedup, version
management, idempotency. Canonical 2026-04-22 pattern: AutoCDC
over hand-rolled
MERGE. - concepts/change-data-capture — third CDC ingest mode on the wiki (after log-based and CDF); the runtime infers deltas from consecutive whole-table snapshots. First-class input mode in AutoCDC.
-
concepts/change-data-capture —
sequence_bycolumn as the declarative primitive for CDC ordering independent of arrival; canonicalises a concern separate from idempotency. -
concepts/probabilistic-data-structure / concepts/probabilistic-data-structure — Databricks' 2026-04-29 sketch-functions launch canonicalises mergeability as the architectural unlock that turns sketches from "a faster percentile" into a storage primitive; every sketch in the launch is mergeable across partitions, time windows, and streaming micro-batches.
- concepts/probabilistic-data-structure / concepts/probabilistic-data-structure /
concepts/probabilistic-data-structure / concepts/probabilistic-data-structure
— the four new first-class SQL / DataFrame / Structured
Streaming aggregates from the 2026-04-29 launch, each
replacing a specific exact-query failure mode (global sort,
full ID shuffle, cluster-wide reduction, composed
GROUP BY). - decision-support-vs-audit-query — the architectural classifier the 2026-04-29 post introduces explicitly to scope where sketch primitives apply. "When to use sketches: Dashboards, trend analysis, monitoring, marketing attribution. When to stay exact: Financial auditing, compliance reporting."
- precomputed-sketch-column-in-delta-table — the intended workflow for the 2026-04-29 launch: build sketches once during ETL, store as BLOB columns in Delta tables, merge on read. Converts a weekly-trending dashboard from a billion-row scan into a 168-sketch merge.
-
set-algebra-on-theta-sketches — audience overlap / incrementality / exclusive reach via union / intersection / difference on kilobyte-scale Theta sketches; microsecond set operations where exact computation requires a cluster-wide shuffle of user IDs.
-
concepts/compute-storage-separation — Lakebase's Pageserver+ Safekeeper vs ephemeral Postgres compute split is a canonical OLTP-shape instance (cf. Aurora DSQL, Snowflake). Second production datapoint via LangGuard (2026-04-27): burst-driven agentic workload exploits the "compute attaches with no data movement" property for scale-to-zero without data cold-start.
- concepts/governed-agent-data-access — runtime control infrastructure for autonomous agent workflows; canonicalised via the LangGuard 2026-04-27 profile. Visibility-gap framing (autonomous-agent logic generated on the fly bypasses conventional SIEM-shape audit).
- concepts/attribute-based-access-control — synchronous allow/deny/modify gate before action execution; the LangGuard control primitive.
- agent-behavioral-baseline — learned baseline of agent behavior from historical trace data, used for anomaly detection; stated roadmap at LangGuard profile time.
- bursty-query-pattern — LangGuard's agentic-workload trace+enforcement traffic extends this concept to the combined write-burst + read-burst shape distinct from the OLAP read-burst framing.
- concepts/scale-to-zero — canonically articulated by LangGuard team as the QRadar-era missing capability for bursty security-telemetry workloads.
- concepts/database-branching — Lakebase's instant copy-on-write branching exploited by LangGuard for governance policy testing — first wiki instance of the primitive at policy-validation altitude.
- concepts/copy-on-write-storage-fork — Lakebase/Neon lineage as second canonical wiki instance after Aurora blue/green; governance-policy-testing as a new use case axis.
- runtime-governance-enforcement-layer — the shape LangGuard canonicalises at agentic-workflow altitude; inline gate on every agent action backed by live knowledge-graph context.
- patterns/database-branch-per-test-over-mocking — clone production trace data via copy-on-write branching in seconds, test new governance policies against real agent behavior in isolation, discard branch.
- agent-provisioned-database — the 2026-04-29 Lakebase launch-partner announcement for Stripe Projects makes Databricks the second launch-side provider in the agent-provisioning protocol. Lakebase/Neon becomes the canonical first- instance of this new concept — a database-tier sibling of agent-provisioned-account with sub-350 ms provisioning + scale-to-zero + copy-on-write branching as the three-pillar substrate contract.
- patterns/partner-managed-service-as-native-binding — Databricks/Neon via Stripe Projects is the third known-use instance of this pattern (after Cloudflare/PlanetScale + Fly.io/Tigris) and the first agent-as-customer instance with a payments-platform orchestrator rather than a compute- platform orchestrator.
- concepts/point-in-time-recovery — canonically disclosed at Lakebase altitude 2026-04-30 with 3.78-second end-to-end recovery for a 32-row deletion incident. On a copy-on- write-capable substrate, PITR collapses into a copy-on-write storage fork at a past timestamp; same operation as branching with a different time parameter.
- concepts/wal-write-ahead-logging — PITR target-times snap backward to the nearest WAL record (12-second snap-back demonstrated on Lakebase POC); a structural property, not a bug, but load-bearing for time-sensitive recovery.
- mock-object-maintenance-cost — the 20-30%-of- test-code maintenance burden Thoughtworks argues cheap Lakebase branching eliminates. Test infrastructure that diverges from production + produces false confidence.
- concepts/integration-tests-against-real-database — the testing discipline cheap branching re-enables: full real-DB behaviour in tests (constraints, transactions, query-planner, lock ordering) rather than mocked surrogates.
- concepts/short-lived-credential-auth — Lakebase's
data-plane auth posture; classic Databricks PATs are rejected,
short-lived OAuth JWTs minted by
databricks postgres generate-database-credentialare required. - branching-is-pitr-with-time-now — architectural
unification canonicalised at Lakebase altitude 2026-04-30:
branching and PITR are the same primitive with different
source_branch_time. Same control-plane call, same storage substrate, same compute-attach step; latency envelopes confirm (1.09 s vs 3.78 s, same order of magnitude). - patterns/database-branch-per-test-over-mocking — CI / QA / IDE workflow pattern that replaces database-interface mocks with per-test / per-PR / per-developer database branches when branching is sub-second + free. Canonical instance: Thoughtworks Backstage POC on Lakebase.
- credential-refresh-cron-as-auth-compat-shim —
pragmatic pattern for bridging short-lived-JWT data planes to
applications expecting long-lived credentials; Thoughtworks POC
used a 50-minute cron rewriting
DATABRICKS_TOKENin.env. - systems/lakehouse-federation — UC's query-federation
surface that exposes external systems as foreign catalogs
governed by UC. First wiki canonical instance (2026-05-15):
Backstage Postgres on Lakebase exposed as
lakebase_bsforeign catalog; standard UC GRANTs replace Postgres native grants, removing the operational↔analytical security-paradigm split. - systems/lakebaseops — open-source Thoughtworks-built
Databricks App for Lakebase DBA automation. Three agents
(Provisioning, Performance, Health) replace 51 historical DBA
tickets; seven scheduled Databricks
Jobs replace pg_cron; monitoring UI surfaces live
pg_statmetrics + slow-query regressions + branch TTL enforcement + 9-KPI adoption dashboard; migration wizard scores ten source engines (Aurora, RDS, Cloud SQL, AlloyDB, Cosmos DB, others) with live AWS+Azure pricing. Inherits governance from UC GRANT - audit trail. Repo:
github.com/suryasai87/lakebase-ops-platform. -
systems/lakebase-mcp — open-source Thoughtworks-built MCP server exposing 46 tools to MCP-capable AI agents (Claude, Copilot, GPT) for Lakebase Postgres access. Dual-layer governance: SQL-statement guard + per-tool access guard across four pre-built profiles (
read_only,analyst,developer,admin) mapping onto the same UC GRANT model; "a coding assistant runs asread_onlyand physically cannot drop a table." Per-statement tool-tag attribution makes "which agent on which branch generated the 4 AM CPU spike?" a one-SQL query. Repo:github.com/suryasai87/lakebase-mcp. -
systems/omnigent — Omnigent (first wiki disclosure 2026-07-23). Open-source AI agent framework with a built-in contextual policy engine that evaluates every tool call against multiple policy layers (intent, risk scoring, PII blocking, custom rules) using deny-wins composition. Introduces intent-based authorization: binds sessions to a declared purpose so out-of-scope actions are blocked even when identity permits them. Targets the confused deputy problem in AI agent systems. Alpha stage.
Recent articles¶
-
2026-10-02 — sources/2026-10-02-databricks-real-time-retail-intelligence-building-e-commerce-recommenda — Real-Time Retail Intelligence: Building E-Commerce Recommendations with Lakebase and AI Search. Reference architecture (Tier-3, include: real-time serving infra, two-path serving design, feature-store/training-serving consistency, cold-start, funnel internals — real system-design substance well >20%) for a production e-commerce recsys, from a real fashion deployment in Asia (1M+ MAU, 100K+ SKUs, ~1,000 events/sec). Everything runs on Databricks under Unity Catalog lineage from raw clickstream to served prediction. Core design: two serving paths split by whether context is known in advance. Path A (homepage/category/email/push — user+surface known ahead) is an async-projected read model: a nightly Workflows job runs the full funnel offline per user and writes a top-N list (50–100/surface) into Lakebase online tables keyed by
user_id + surface; serving = plain KV lookup ("no model inference, no vector search, no feature assembly"). Path B (similar-items / complete-the-look / re-ranked search — context exists only at request time) runs the full three-stage retrieval → ranking funnel synchronously inside one custom MLflow PyFunc endpoint (co-located inference): Stage 1 blended user⊕session query embedding → AI Search hybrid retrieval = ANN + hard metadata filters (out-of-stock/ineligible never enter candidate set) → 200–500 candidates in one round-trip; Stage 2 Lakebase point-lookup features + cross features → LightGBM conversion scoring (CPU, no GPU); Stage 3 config-driven business-rule re-rank → 10–20 items. In-session signals arrive in the request payload and bypass the lakehouse on the hot path; Zerobus ingests clickstream only for offline training. Explicit latency-budget fallback: Path B → cached popular items or Path A on overrun. Cold start handled on both axes (demographic default embedding for new users; attribute+visual embedding + nearest-neighbor score inheritance for new SKUs). Feature store gives training-serving consistency; weekly champion/challenger retrain with drift detection, request-level serving↔training correlation, and position-aware training. No new pages — taxonomy gate applied: two-path serving maps onto async-projected-read-model (Path A) + co-located-inference + retrieval-ranking-funnel (Path B) + cheap-approximator-with-expensive-fallback (fallback); champion/challenger recorded as tags+prose (single source, per "tag don't mint"); AI Search reuses systems/mosaic-ai-vector-search (confirmed rebrand). Updated systems systems/lakebase, systems/mosaic-ai-vector-search, systems/zerobus-ingest, systems/databricks-model-serving, systems/mlflow, systems/lightgbm, systems/unity-catalog; concepts concepts/medallion-architecture, concepts/retrieval-ranking-funnel, concepts/cold-start, concepts/feature-store, concepts/training-serving-boundary, concepts/hybrid-search; patterns patterns/async-projected-read-model, patterns/co-located-inference-in-serving-layer, patterns/cheap-approximator-with-expensive-fallback. ~18 wiki pages touched. URL verbatim from raw frontmatter. Caveats: vendor reference architecture from one unnamed customer; latencies qualitative ("<2-digit ms"); 10–30% conversion-uplift is an industry benchmark, not this system's measured result. -
2026-10-01 — sources/2026-10-01-databricks-lakebase-postgres-branch-based-restores-for-fast-recovery-at-scale — Lakebase Postgres branch-based restores for fast recovery at scale. (Tier-3, include: three-stage traditional-PITR breakdown + compute/storage split + immutable-timeline storage model + branch-at-timestamp restore mechanism, architecture well >20% of body.) Reframes point-in-time recovery on Lakebase: traditional managed OLTP (RDS) restore is a data-movement job that scales with data size — provision a coupled compute+storage instance, hydrate the snapshot from S3 onto the Postgres disk, then replay WAL to the target time; "unless the database is small, PITR is almost always a multi-hour operation." An HA replica is no substitute (the bad write "is already on the standby"). On Lakebase, because the Safekeeper + Pageserver + object-storage tier keeps history as an immutable addressable timeline (old page versions never overwritten), a restore is a branch at a point in history — the control plane maps a timestamp → LSN, creates a branch, attaches compute. The branch "points at the image and delta layers that already exist" (COW fork) so "a restore is metadata: a pointer to a point in history" — no copy, no replay. Load-bearing claim: restore time is independent of data size ("restoring 100 TB is as fast as restoring 10 GB"), improving RTO at scale; only validation is variable-cost. Cited pain survey (50 devs, 1 TB+ prod Postgres): 30% down 3+ hours, only 21% recovered <60 min. Also frames fast branch-at-timestamp restore as an agent/product undo primitive (Replit, v0: checkpoint = snapshot of
main+ stored ID next to code version; undo = restore onto the live branch). No new pages — pure canonicalization into existing PITR/branching/COW/compute-storage-separation vocabulary. Created (1): source page. Updated (8): concepts concepts/point-in-time-recovery (size-independent-restore Seen-in), concepts/database-branching (branch-as-restore Seen-in), concepts/compute-storage-separation (restore-as-metadata Seen-in), concepts/copy-on-write-storage-fork (image+delta-layers Seen-in), concepts/immutable-object-storage (addressable-timeline Seen-in), concepts/rpo-rto (size-independent-RTO Seen-in); systems systems/lakebase (new Branch-based restores section + Seen-in), systems/pageserver-safekeeper (restore-path role Seen-in), systems/aws-rds (copy-and-replay-foil Seen-in); companies/databricks (this entry). index.md Latest-additions top entry hand-updated; category/landing/sources indexes are build-time auto-generated. ~11 wiki pages touched. URL verbatim from raw frontmatter (https://www.databricks.com/blog/lakebase-postgres-branch-based-restores-fast-recovery-scale; raw filename slug truncated...for-fast-recovery-at-e79eb7ef). Flipped raw frontmatteringested: false→ingested: true. -
2026-09-30 — sources/2026-09-30-databricks-a-practical-guide-to-cost-optimization-with-lakebase-postgre — A practical guide to cost optimization with Lakebase Postgres. Cost-optimization guide (Tier-3, include: storage/compute-separation economics, COW branching, CDC-based reverse-ETL, working-set sizing,
max_connectionsrule — real system internals well >20%) for Lakebase. Thesis: separated storage/compute + serverless compute make normally-expensive operations cheap by design — branches, read replicas, and HA each add compute that reads the shared storage layer, duplicating no storage; autoscaling + scale-to-zero bills compute to demand. Discloses pricing/sizing specifics: default new project 8–16 CU autoscale, scale-to-zero after 24h; always-on pricing = 25% discount on baseline when scale-to-zero off (covers HA compute too); compute cache up to 75% of RAM;max_connections= min(max CU, 8 × min CU) with a 10k-client connection pooler (systems/pgbouncer) as the cheaper alternative to oversizing (concepts/connection-pool-exhaustion); PITR window 2–30 days; storage snapshot $0.090/GB-mo (~74% cheaper), PITR $0.200/GB-mo (~42% cheaper). Operational levers: sync only the working-set subset of a Delta table via a rolling MV + its automatic change data feed (a CDC-driven reverse-ETL pattern); match sync mode to freshness (Snapshot >10% churn, up to 10× cheaper than Triggered; Triggered = incremental sweet spot; Continuous = highest cost); bin-pack multiple tables into one SDP pipeline; size compute to the working set, not disk size. Three metered buckets (database compute, storage, Synced-Table pipeline compute) all insystem.billing.usage, storage split byproduct_features.lakebase.storage_type(BRANCH_DATA_STORAGE/BRANCH_CHANGE_STORAGE/BRANCH_HISTORY_STORAGE). No new pages — pure canonicalization into existing Lakebase/concept vocabulary. Enriched systems/lakebase (cost-model Seen-in), systems/lakeflow-spark-declarative-pipelines + systems/unity-catalog (Synced-Tables Seen-in), and concepts concepts/compute-storage-separation, concepts/scale-to-zero, concepts/working-set-memory, concepts/connection-pool-exhaustion, concepts/point-in-time-recovery, concepts/database-branching, concepts/change-data-capture, concepts/elt-vs-etl, concepts/materialized-view, concepts/bin-packing (Seen-in). URL verbatim from raw frontmatter (practical-guide-cost-optimization-lakebase-postgres; raw filename slug truncated...lakebase-postgre). Caveats: list-price only (no contractual discounts); absolute branch-storage $/GB not stated; SDK/DABs sizing snippets referenced but not reproduced. -
2026-09-29 — sources/2026-09-29-databricks-your-data-your-storage-your-rules-a-2026-guide-to-storing-un — Your data, your storage, your rules: a 2026 guide to storing Unity Catalog managed tables. Storage-model/positioning guide (Tier-3, borderline-include: storage-ownership + placement-hierarchy design is genuine reusable system-design substance, >20% of body) on where UC managed table data physically lives. Load-bearing claim: managed tables get catalog-owned automation (layout/tuning/cleanup/predictive optimization/commit coordination) while the files stay in cloud storage the customer owns — S3/ADLS/GCS in the customer's own account — not provider-controlled/proprietary storage. This is the control-plane / data-plane split applied to a catalog: UC governs, the customer account holds the bytes ("files remain accessible in your cloud account while Unity Catalog governs access to them"). Placement = a managed storage location inherited metastore → catalog → schema, most-specific-wins, living inside a UC external location (the same governed cloud-path+credential object that governs external tables); mutable + non-retroactive (
ALTER CATALOG/SCHEMA … SET MANAGED LOCATIONredirects new tables, existing stay put). Per-catalog/schema physical placement is a concrete GDPR-segregation / regional-residency / cost-allocation boundary lever (logical org + ABAC alone "satisfies standard GDPR data segregation requirements"; physical separation for boundaries that must reach the storage).ALTER TABLE … SET MANAGEDconverts external→managed by copying data + transaction log into the resolved location. Openness (Iceberg/Delta + Iceberg REST Catalog + UC open APIs + credential vending, engines: Spark/Trino/Flink/Kafka Connect/Snowflake) framed as not lock-in. No new pages — pure canonicalization into existing vocabulary; customer-owned-storage / bring-your-own-storage recorded as tags + prose (single-source; per the taxonomy gate → tag, don't mint). Enriched systems/uc-managed-tables (new storage-ownership + placement model section), concepts/control-plane-data-plane-separation (catalog-govern-customer-bucket Seen-in), concepts/data-residency (per-catalog/schema physical-placement Seen-in), concepts/open-table-format (anti-lock-in Seen-in), systems/unity-catalog / systems/uc-credential-vending / systems/iceberg-rest-catalog-scan-api (Seen-in). URL verbatim from raw frontmatter (your-data-your-storage-your-rules-2026-guide-storing-unity-catalog-managed-tables; raw filename slug truncated...-storing-un-ee2e9e6d). Caveats: positioning guide — no wire-protocol / commit-coordination / layout internals / latency/throughput numbers; "satisfies GDPR" is vendor framing, not legal advice. -
2026-09-28 — sources/2026-09-28-databricks-lakebase-search — Lakebase Search: State-of-the-art full text and vector search for Postgres. Launch/architecture post (Tier-3, include: distributed-systems internals — vector-index architecture, storage/compute separation, quantization, scaling trade-offs — well >20%) shipping two GA extensions on Lakebase Postgres (AWS + Azure): lakebase_vector (ANN vector search) and lakebase_text (BM25 full-text). Core argument: pgvector hits a wall at scale on three structural axes — cost scales with data volume not usage (its HNSW index must be RAM-resident: 768-dim float32 ≈ 3.3 KB → 100M rows ≈ 330 GB RAM, disk-spill = 10×–50× slowdown, no working-set notion); index maintenance blocks the DB (~50 h builds, table-locking
REINDEX); and a single query can't parallelize the HNSW scan. lakebase_vector rebuilds the index around Lakebase's storage/compute separation: hierarchical IVF clustering (contiguous cluster blocks, in-memory centroid scoring → hundreds of random hops become a handful of large sequential reads) + binary RaBitQ quantization (arXiv 2405.12497, ~1 bit/dim, ~32×) over object storage — fast both hot-in-RAM and cold-on-object-storage, with a scan-codes-then-rerank funnel. Yields stateless scale-to-zero (P90 cold start 1.13 s; 100M vectors on 1 CU), parallel/Spark-offloadable (LTAP) builds down to minutes (roadmap), per-query core parallelism, and inline SQL predicate filtering. VectorDBBench LAION 100M: 2× throughput, 4× cheaper vs cloud pgvector, P99 71 ms @ 97% recall. lakebase_text adds corpus-aware BM25 (global IDF, beatstsvector+GIN via posting-block upper-bound pruning); combined they give native hybrid search — vector + BM25 + SQL filters + joins in one query (patterns/parallel-retrieval-fusion). Conexiom: BM25 hybrid over 100M+ rows at ½ compute footprint, 3× lower spend, 5× throughput. New pages (6): systems systems/lakebase-vector, systems/lakebase-text, systems/rabitq, systems/vectordbbench; concepts concepts/hybrid-search (≥2 sources — also Dropbox Dash, Cloudflare AI Search, Expedia), concepts/ivf-inverted-file-index. Enriched systems/lakebase (Lakebase Search section), systems/hnsw (pgvector-wall Seen-in), systems/diskann + concepts/vector-similarity-search + concepts/quantization (binary/RaBitQ 1-bit) + concepts/compute-storage-separation (search-index axis) + concepts/scale-to-zero (vector-search cold start) + systems/bm25 (native-Postgres-BM25) + systems/mosaic-ai-vector-search (positioning contrast). URL verbatim from raw frontmatter (lakebase-search-state-art-full-text-and-vector-search-postgres; raw filename slug expanded). Caveats: vendor marketing post; benchmark numbers are Databricks' own (pgvector/DiskANN tested only on a single large instance — concepts/benchmark-methodology-bias); Spark-offloaded minutes-scale builds are "stay tuned" roadmap; IVF fan-out / rerank shortlist / filtered-recall degradation undisclosed; Databricks positions managed Databricks AI Search (systems/mosaic-ai-vector-search) as the tuning-free alternative. -
2026-09-28 — sources/2026-09-28-databricks-how-databricks-rolls-out-frontier-models-to-12000-employees — How Databricks rolls out frontier models to 12,000 employees on Day 1. Internal-practice post (Tier-3, borderline-include: model release lifecycle + four-budget architecture + stratified cost-comparison methodology well >20%) on the playbook that gives 12,000+ employees Day-1 access to newly-released models while assessing whether they're genuine long-term workhorses. Motivating failures: "frontier" models often aren't (Opus 5.0 ranked lower than 4.8 while costing more — concepts/efficiency-frontier) and naive rollout explodes cost (a no-mitigation "GPT Astra" control group spent +60% overnight). The three-stage model release lifecycle (patterns/experimental-tier-model-promotion — new pattern page): (1) immediate experimental access to all employees via Unity Gateway server-side designation + UG CLI laptop config push over MDM into Claude Code / Codex / Omnigent (models shown with an "Experimental" tag); (2) budget-constrained exposure — the four-budget architecture (monthly max, daily runaway, [NEW] quality-frontier budget rationing premium models at a 2–3× premium, [NEW] experimental budget bounding new-model exposure — systems/unity-ai-gateway-budgets); (3) promote or drop on three signals — private benchmarks (offline + online side-by-side PR creation, concepts/llm-as-judge, concepts/benchmark-methodology-bias), user reports, and OTel cost tracking. Cost comparison stratifies sessions (single/multi-turn × file-edits) and re-weights before comparing (sibling of patterns/stratified-evaluation-sampling) → Opus 4.8→5.5 −29% $/session, GPT-5.6→6 Sol −48%. Outcomes: Opus 5 / GPT-6 Sol / Luna → GA by Day 3; Opus 5.5 slated as Claude Code default; GPT-6 Sol not the Codex default but added to the smart router's toolkit (concepts/model-first-routing). New pages (2): patterns/experimental-tier-model-promotion (≥2 sources — this + 2026-08-07 cost-management post) + the source page. Taxonomy gate applied — NO other new pages: session-cost-normalization / stratified-cost-comparison, four-budget-architecture, day-1-model-access recorded as tags + prose (map into concepts/efficiency-frontier, patterns/stratified-evaluation-sampling, systems/unity-ai-gateway-budgets per "when unsure, tag — don't mint"). Enriched systems/unity-ai-gateway (Day-1-rollout-plane + UG-CLI-push section), systems/unity-ai-gateway-budgets (four-budget-architecture section), concepts/efficiency-frontier (promote/drop-gate Seen-in), concepts/model-first-routing (promote-to-router-toolkit variant), systems/omnigent / systems/claude-code / systems/codex-cli / systems/opentelemetry (Seen-in). URL verbatim from raw frontmatter (
how-databricks-rolls-out-frontier-models-12000-employees-day-1). Caveats: product-adjacent; model names (Opus 5.5, GPT-6 Sol, GPT Astra, Claude Fable, Luna) are forward-dated/placeholder — mechanisms are the durable content; all costs relative ($/session deltas, no absolute spend); router internals documented elsewhere (2026-08-13). -
2026-09-25 — sources/2026-09-25-databricks-from-data-to-dialogue-how-sp-global-energy-made-its-structured-data-estate-conversational — From Data to Dialogue: How S&P Global Energy Made Its Structured Data Estate Conversational with Databricks Genie Agents and MCP. Customer story (Tier-3, borderline-include: real three-layer architecture + MCP composition + governance design well >20%) on making S&P Global Energy's structured-data estate (Chemicals, Crude Oil, Refined Products, Gas & Power, LNG, …; some across non-Databricks sources) conversational for external consumption by AI agents via MCP. Three-layer design, each owned by the right people: (1) SMEs curate one Genie Agent per dataset group — no agent code — native tables via Unity Catalog, non-Databricks tables via Lakehouse Federation (no ETL), enriched with descriptions, example queries, trusted assets, and business definitions (e.g. "floating storage = cargoes idling ≥3 days below a threshold speed") — the patterns/specialized-agent-decomposition shape ("resist… one agent to rule them all"); (2) every Genie Agent is automatically a Databricks-managed MCP server at
/api/2.0/mcp/genie/{genie_space_id}— nothing to deploy — with a two-tool ask-then-poll surface (genie_query_space+genie_poll_response) running async on a SQL warehouse, UC-governed and platform-authenticated ("we inherited the [security layer] we already had"); (3) FastMCP proxy composes group Genie MCP servers into per-commodity composite endpoints with name-spaced tools (patterns/mcp-as-centralized-integration-proxy) so an agent hits one endpoint per commodity and its LLM routes or fans-out cross-group questions. One MCP-standard bridge serves internal agents, customer-facing AI, and external customers' own MCP agents. Genie Agent Benchmarks give a continuous curate→benchmark→improve loop; best adoption signal = "how often SMEs agreed with Genie's generated SQL" ("measure trust, not just latency"). Outcome: new conversational data domains live in days, not development cycles; SMEs became publishers, not requesters. New page: systems/fastmcp (proper-noun composition framework, taxonomy-exempt). No new concepts/patterns minted — genie-agent-benchmarks, ask-then-poll, and external-consumer MCP recorded as tags + prose (single source; per the taxonomy gate → reuse concepts/text-to-sql, concepts/governed-agent-data-access, patterns/specialized-agent-decomposition, patterns/mcp-as-centralized-integration-proxy). Enriched systems/databricks-genie (Genie-Agent-managed-MCP + benchmarks section), systems/model-context-protocol (managed-MCP-per-Genie-space Seen-in), systems/unity-catalog (governs-Genie-MCP+federated Seen-in), systems/lakehouse-federation (Genie-substrate Seen-in), systems/databricks-sql-warehouses (async-Genie-execution Seen-in), concepts/text-to-sql (SME-curated-semantic-layer Seen-in), concepts/governed-agent-data-access (external-consumer Seen-in), patterns/mcp-as-centralized-integration-proxy (FastMCP-composition Seen-in), patterns/specialized-agent-decomposition (one-Genie-per-dataset-group Seen-in). URL verbatim from raw frontmatter (data-dialogue-how-sp-global-energy-…; raw filename slug truncatedstructur). Caveats: customer-authored; no hard numbers (latency/cost/QPS/accuracy); FastMCP proxy auth-passthrough internals and federation enforcement mechanism undisclosed. -
2026-09-24 — sources/2026-09-24-databricks-how-i-built-agent-based-security-reviews-on-databricks — How I built agent-based security reviews on Databricks. First-person account (Tier-3, borderline-include: real multi-agent system architecture + governance + HITL design >20%) of extending an existing security-review process with an agent-based layer built entirely on the Databricks platform. Core design choice: not one general "security reviewer" agent but seven focused agents (patterns/specialized-agent-decomposition) — Intake, Risk Assessment, Requirements, Specialized Review (browser-extension threat modeling, vendor assessment), Validation, Workflow, Learning — each with bounded responsibility so "its behavior remains inspectable and testable, and a change to one does not silently affect another." Platform: Unity Catalog as the governed system of record (standards / requests / evidence / model outputs / decisions under one permission + lineage model → traceable dashboard metrics, concepts/audit-trail); Databricks-hosted Claude Haiku → Sonnet → Opus via the Foundation Model API as the reasoning layer, tiered by task difficulty (patterns/cheap-approximator-with-expensive-fallback, concepts/model-first-routing); Lakeflow Jobs orchestrating the agents on serverless compute; two Databricks Apps (conversational intake front door + consultation mode; executive dashboard). Trust thesis (not model quality): automated completion confined to predefined low-risk request classes, backed by acceptable evidence ("an assertion with no evidence is treated as missing information"), with a conservative default on uncertainty — "When evidence is missing or contradictory, the system does not guess. It defaults to the more conservative risk tier... It does not find its way to approval" (fail-closed on the security invariant, concepts/human-in-the-loop escalation for novel/high-risk/ambiguous). A Learning agent compares reviewer edits to original output to improve prompts + standards (humans apply the change; agents don't mutate prod). No new pages — pure canonicalization into existing vocabulary. Enriched patterns/specialized-agent-decomposition (security-review-intake instance), concepts/human-in-the-loop (evidence-gated escape valve), concepts/fail-open-vs-fail-closed (conservative-default agentic decisioning), concepts/separation-of-concerns (bounded agents), systems/unity-catalog (agent system-of-record), systems/databricks-apps (seventh face: intake + dashboard), systems/lakeflow-jobs (agent orchestrator), systems/databricks-foundation-model-api (first Claude-hosting + task tiering), patterns/cheap-approximator-with-expensive-fallback (static task tiering), patterns/human-calibrated-llm-labeling (reviewer-edit feedback), concepts/least-privileged-access (bounded agent authority), concepts/audit-trail, concepts/retrieval-augmented-generation (grounding in standards), concepts/structured-output-reliability (checkable requirements), concepts/model-first-routing (task-difficulty variant). URL verbatim from raw frontmatter. Caveats: vendor first-person account; dashboard numbers redacted; inter-agent coordination protocol, eligibility-criteria representation, and risk-tier encoding undisclosed; single "days → minutes" outcome is illustrative.
-
2026-09-23 — sources/2026-09-23-databricks-how-concurrence-governs-clinical-ai-at-a-trillion-token-scal — How Concurrence governs clinical AI at a trillion-token scale with Unity Gateway. Customer story (Tier-3, borderline-include: event-sourcing + compliance-gated routing + production numbers >20%) on Concurrence, a healthcare company running clinical AI agents at ~1.2T annualized input tokens (~100.8B tokens / 11.2M LLM calls per 30 days, ~5× monthly growth). Three load-bearing ideas: (1) an event-sourced "world model" — new patient info recorded as immutable events rather than overwriting records, current state computed from history with preserved provenance (concepts/log-as-truth-database-as-cache); this makes replay-based agent testing free and side-effect-safe (patterns/snapshot-replay-agent-evaluation, sim+eval ~7× production traffic). (2) A shared Databricks foundation: events → Zerobus → governed Delta (2.7M world-model events/mo, 90k/day peak), Spark Declarative Pipelines derive the world model, Unity Catalog governs per-tenant schema + service principal (concepts/tenant-isolation), Lakebase serves operational state (retiring a homegrown prompt-log store + reverse-ETL for Synced Tables), apps on Databricks Apps. (3) Compliance-gated model routing via Unity Gateway — "routing is compliance-gated before it is cost-gated"; batch
ai_queryon BAA-covered Databricks-hosted Claude, endpoint resolver restricted to the BAA-covered namespace (structurally blocks PHI to uncovered models, also protects self-harm/suicidal-ideation/medical-emergency classification); real-time inference built + feature-flagged behind a synthetic canary pending compliance coverage. Coding agents fully on Unity Gateway'sugCLI (patterns/on-behalf-of-agent-authorization; July: 14 users, 35.85B tokens, 95.37% cache reads; ~360k requests + ~61B tokens since Jul 10; billing viasystem.ai_gateway.usage). 14 production-inference models, 46-model governed catalog; testing Smart Routing + clinical reasoning benchmarks; exploring Omnigent. New page: systems/concurrence (proper-noun customer platform). Enriched systems/unity-ai-gateway (compliance-gated routing +ug-at-scale), concepts/centralized-ai-governance (compliance-as-ordering-constraint), concepts/model-first-routing (compliance-gated variant +compliance-gated-routingalias), concepts/log-as-truth-database-as-cache (clinical world-model instance), systems/zerobus-ingest / systems/lakeflow-spark-declarative-pipelines / systems/lakebase / systems/databricks-apps / systems/unity-catalog (Concurrence deployment faces), patterns/snapshot-replay-agent-evaluation + patterns/telemetry-to-lakehouse.compliance-gated-routingrecorded as tag + prose + alias, not minted as a standalone page (1 source; per the taxonomy gate — reuse concepts/model-first-routing + concepts/centralized-ai-governance). URL verbatim from raw frontmatter. Caveats: customer story, light on mechanism (no world-model schema, resolver/router internals, Lakebase sizing, or latency); real-time path not yet on Unity Gateway. -
2026-09-22 — sources/2026-09-22-databricks-genie-one-mcp-give-any-ai-agent-the-right-business-context — Genie One MCP: Give any AI Agent the Right Business Context. Launch of an MCP server exposing Genie One's governed conversational analytics to any MCP-compatible agent (ChatGPT, Claude, Copilot, Claude Code). Thesis: direct data access ≠ business context — direct source wiring gives access but not shared governed meaning, introducing four gaps (accuracy / cost & latency / governance / consistency); route every agent through one governed layer backed by Genie Ontology instead. Concrete architecture: five-tool contract (
genie_ask→conversation_id+response_id;genie_poll_response→ progress + answer + Explore-in-Databricks deep link;genie_get_query_result;genie_cancel_response;view_askas the MCP Apps interactive-view variant) +warehouse_id_metapin. Identity: OBO recommended (end-user OAuth token → UC privileges + row filters + column masks in the user's context; two users, same question, appropriately scoped answers, no per-user prompt logic); M2M flattens identity and is the named anti-shape; external MCP connections are UC securables. Management: tune Genie in Databricks not from the client system prompt; prefer the scoped Genie Agent MCP server (/api/2.0/mcp/genie/{genie_space_id}) for one curated domain; 90-second SQL timeout + workspace Genie QPM limit; one OAuth app per client platform; validate governance by impersonation. New page: systems/databricks-genie-one (proper noun — Genie One + Genie One MCP). Extended systems/databricks-genie-ontology (Genie-One-MCP-exposure Seen-in + frontmatter), systems/databricks-genie (Genie One related-link), systems/model-context-protocol (MCP-Apps + governed-analytics-surface + external-MCP-as-UC-securable Seen-in), concepts/governed-agent-data-access (four-failure-mode Seen-in + frontmatter), patterns/on-behalf-of-agent-authorization (OBO-for-analytics-MCP Seen-in + M2M-anti-shape + frontmatter), patterns/mcp-as-centralized-integration-proxy (governed-intent-surface Seen-in + frontmatter). Tier-3 vendor launch post, borderline-include (five-tool contract + OBO/U2M/M2M identity model + MCP Apps + operational constraints well above the 20% architecture bar). URL verbatim from raw frontmatter. Caveats: no scale numbers; ontology internals deferred to the 2026-09-15 source; OBO enforcement requires the user token to actually flow through. -
2026-09-18 — sources/2026-09-18-databricks-database-branching-a-developers-guide-to-git-style-workflows — Database Branching: A Developer's Guide to Git-Style Workflows. Developer-education explainer (Tier-3, borderline-include: copy-on-write mechanics + safety architecture >20%) that gives the platform-neutral canonical framing of database branching — an isolated DB environment forked from a parent's schema+data at a point in time, DB-tier analog of a
git branch, made practical by copy-on-write (branch shares parent pages, stores only divergent data). Two canonicalisations beyond prior branching-series entries: (1) the workflow divergence from Git — "you usually don't merge changes from a database branch back into the parent database. Instead, migration files remain the durable source of truth" — you test the migration on the branch against realistic data, then the pipeline applies that same migration to the target (version-controlled migrations stay the durable artifact); (2) AI agents make branching infrastructure-critical — "hundreds or thousands of short-lived environments" per agent fleet where full copies are too slow/expensive, and branching reduces the blast radius of agent mistakes by handing a branch instead of production write access. Three workflow archetypes: production-like baselines (anALTER TABLE … ADD COLUMN … NOT NULLpasses on empty DB, fails on millions of rows → the realistic branch catches it), per-PR isolation (patterns/database-branch-per-test-over-mocking), discard-and-refork failure recovery (PITR-adjacent). Six operational-safety guardrails: protect parent branches, safe/mock data on ephemeral branches, TTL, migrations as source of truth, reproducible-not-repaired branches, least-privilege + Unity Catalog access governance. Worked CoW cost model: 40 GB parent branched twice = +80 GB full-copy vs ~5.6 MB CoW (1.6 MB + 4 MB divergence). No new pages — pure canonicalization; maps entirely into existing vocabulary. Extended: concepts/database-branching (workflow-divergence + agent-scale legs), concepts/copy-on-write-storage-fork (worked cost model + symmetric-fork + agent-fleet affordability), concepts/evolutionary-database-design (migrations-as-source-of-truth), patterns/database-branch-per-test-over-mocking (per-PR CI shape), patterns/per-developer-database-branch-paired-with-code-branch, systems/lakebase (CoW branching substrate). Caveats: Tier-3 vendor explainer funneling to Lakebase + a Postgres tutorial; no new architecture beyond the LangGuard/Stripe/Backstage/evolutionary-DB three-parter; the 1.6 MB/4 MB divergence numbers are illustrative; "branches aren't merged back" is normative, not a hard constraint. -
2026-09-15 — sources/2026-09-15-databricks-data-ontology-defined-the-context-layer-your-ai-agents-are-missing — Data Ontology defined: The context layer your AI agents are missing. Interview (Richard Tomlinson) defining a data ontology as the semantic context layer that captures what data means (definitions, relationships, calculations, authoritative sources, expertise, permissions) vs a schema, which only captures structure. Thesis: analysts historically supplied business context as tribal knowledge; AI agents have no human in the loop, so missing context is filled with fluent-but-fabricated inference (the "24 customers" board-briefing anecdote). Semantic layers / knowledge graphs became shelfware because teams tried to model the entire enterprise manually; the remedy is "model the head, learn the tail" — govern the small set of concepts that cannot be wrong (revenue, compliance, core KPIs), continuously learn the long tail from existing dashboards/queries/notebooks/usage, ranking by authority. Governance has two jobs: permission-at-retrieval (the ontology must not be a back door around ACLs — enforced via Unity Catalog, so two users get different-but-correct answers; concepts/governed-agent-data-access) and trust/authority ranking (certification, lineage, usage, provenance teach the agent what to believe). Implementation: Genie Ontology, the automatic context layer under Genie One / Genie Agents — internal benchmark 84.5% first-attempt vs 52.4% for the strongest general-purpose coding agent (28 enterprise questions), ~2× faster. New pages: concepts/data-ontology (≥2 sources — this + the 2026-09-03 "Governance beyond security: knowledge, context & ontology" post) and systems/databricks-genie-ontology (proper noun). Enriched concepts/llm-hallucination (enterprise-plausible-wrong-answer + fabricated-number), concepts/governed-agent-data-access (permission-at-retrieval + authority-ranking), concepts/knowledge-graph (ontology-vs-schema framing), systems/databricks-genie (named Genie Ontology section). "Model the head / learn the tail" recorded as prose + tag, not minted as a pattern (pending a clearer second independent source per the taxonomy gate). Tier-3 vendor thought-leadership / launch-adjacent interview but architecture content (ontology-vs-schema, permission-at-retrieval, model-head/learn-tail, authority ranking, benchmark) is >20%; benchmark is internal (28 Qs, unnamed baseline) and the "learn the tail" inference mechanism is undisclosed.
-
2026-09-19 — sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection — RADAR: Catch gray failures with anomaly detection. Databricks' internal reliability-monitoring system RADAR (Reliability Anomaly Detection, Alerting, and Root-cause analysis) built to catch gray failures — partial, silent outages where every dashboard reads green while a slice of customers fails (the framing is differential observability, Huang et al. HotOS 2017). Motivating case: a deploy bug fails ~1-in-20 credit-card checkouts for ~6.5 h, each returning a plausible
INVALID_ARGUMENT, invisible to aggregate thresholds. Four-stage pipeline: (1) reliability metrics — record error count and distinct-user count per error code × region (Zerobus ingest → Delta, UC + Metric Views governance; patterns/telemetry-to-lakehouse); (2) anomaly detection — unsupervised streaming SPOT (Streaming Peaks-Over-Threshold, Extreme Value Theory; Siffer et al. KDD 2017), 14-day normal window, single risk parameter instead of per-metric thresholds (MLflow train → Model Serving → Workflows); (3) alerting — enrich → filter → dedupe → routed ticket (Databricks SQL Alerts), the alert-fatigue defense for many-series detection; (4) root-cause analysis — ticket carries deep-dive + AI/BI Genie-backed dashboard link (AI/BI Dashboards). Ships as one Declarative Asset Bundle (DAB) + a public GitHub scaffold an AI agent builds from a prompt. Metric-agnostic (billing, conversion, model/data drift). Results (vendor-reported): 95% reduction in incident-discovery time at >90% precision, no human needed to spot the pattern. New pages: concepts/anomaly-detection (SPOT/EVT primitive; ≥2 sources — also touched by grey-failure + Airbnb backtesting) and systems/databricks-radar (proper noun). Enriched concepts/grey-failure (differential-observability + user-error-spike detector + checkout example), concepts/observability, concepts/alert-fatigue (enrich/filter/dedupe fix), systems/zerobus-ingest (RADAR ingest layer). Open-loop (routes to a human, not auto-remediation → contrast with patterns/closed-loop-remediation, not an instance). Tier-3 vendor post, product-adjacent (build-it-yourself CTA) but architecture content well above 20%; headline numbers self-reported/unbaselined; SPOT internals + multi-series correlation undisclosed. -
2026-09-14 — sources/2026-09-14-databricks-managed-postgres-what-lakebase-actually-takes-off-your-plate — Managed Postgres: What Lakebase Actually Takes Off Your Plate. Positioning/overview post (Tier-3, borderline-include: operational/architecture content >20%) that frames "managed Postgres" as a spectrum of operational ownership, not a binary, and proposes a four-axis test — maintenance/patching, scaling, HA/failover, backups & recovery — then maps Lakebase against it. Durable systems point: in-region failover is replacement of stateless compute, not replica promotion ("replace the failed compute outright since it holds no durable local state"), because durable pages+WAL live in Pageserver/Safekeeper under storage/compute separation; secondary compute runs in separate AZs, auto-promoted, endpoint unchanged. Failover quality is the RPO/RTO test ("some lose seconds of writes … others lose none"; no defined numbers = "a guess"). Scaling = serverless + scale-to-zero (up to 5× writes vs stock Postgres, vendor-reported); backups = PITR with configurable 2–30 day history + scheduled snapshots. The one candid gap: cross-region disaster recovery is Private Preview, AWS-only, manual failover, customer-managed recovery. AI axis rests on pgvector (vector search + LLM/agent memory co-located with operational data, avoiding the sync tax); dev-experience axes = built-in PgBouncer (pooling) + copy-on-write branching for evolutionary DB dev; migration guidance names logical replication for near-zero-downtime cutover; security = CMK/envelope encryption, Unity Catalog ABAC, default-on audit logging. No new pages — pure canonicalization; the article maps entirely into existing vocabulary. Extended: systems/lakebase (new managed-Postgres coverage-scorecard section), concepts/rpo-rto, concepts/serverless-compute, concepts/point-in-time-recovery. Caveats: marketing framing, 5× and coverage claims are Databricks' own/unaudited; no new architecture disclosures beyond prior Lakebase posts; cross-region DR genuinely not yet automatic.
-
2026-09-10 — sources/2026-09-10-databricks-improving-lakebase-postgres-compute-cache — Improving Lakebase Postgres compute cache. How Lakebase reworks the DRAM/NVMe cache between the query engine and disaggregated object storage. Under disaggregated storage Lakebase reads bypass the OS filesystem, so the classic Postgres shared buffers + OS page cache scheme both wastes RAM (double buffering — 2 GB to cache 1 GB) and loses its kernel cache tier. Interim answer was a two-tier compute cache: in-memory shared buffers + an autoscaling NVMe local file cache (LFC), with shared buffers capped at 1 GB so min-CU computes didn't over-consume memory — which forced most hits into the slower LFC on large working sets (a miss across both tiers routes a
GetPageto storage). Part-1 fix (live today, fixed computes CU ≥ 80): disable LFC, setshared_buffers = 75% of DRAM(80 CU →15278640≈ 116 GiB). But Postgres's process-per-connection model makes big shared buffers impossible without huge pages: 32 GB buffers × 512 backends ≈ 4.3 B 4 KB PTEs ≈ 32 GB of page tables, dwarfing the TLB. Lakebase backs large computes with explicit 2 MB HugeTLB pages (not THP), consistent across host → hypervisor → guest (it runs in guest VMs on bare metal, so translation crosses two virtualized layers), pre-provisioned at VM init with surplus released at startup — benchmarks: ~40% lower tail read latency, ~30% lower CPU (show huge_pages→on). Production rollout (Aug 2026, three large endpoints): ~2× / ~1.3× / ~2× throughput, storageGetPage/s~8K → ~1.5K (≈5× fewer reads), compute cache hit rate ≈ 100%, one workload's CPU 20 → 4 cores. Part 2 (roadmap): dynamic shared buffers + huge pages scaled in concert on autoscaling computes, contributed upstream to Postgres. No new pages — maps entirely into existing vocabulary; double-buffering / huge-pages / TLB-overhead recorded as tags + prose per the taxonomy gate, not minted. Extended: systems/lakebase (new compute-cache section), systems/postgresql, systems/pageserver-safekeeper, concepts/compute-storage-separation, concepts/working-set-memory, concepts/cache-hit-rate. Direct follow-on to the 2026-08-31 autoscaling post. Caveats: large shared buffers live only for fixed CU ≥ 80 (autoscaling still on LFC);15278640/onare 80-CU-specific; benchmark %s are Databricks' own; LFC retired-in-current-form, not deleted. -
2026-09-09 — sources/2026-09-09-databricks-evaluation-first-ai-agents-how-zepto-scales-customer-support — Evaluation-First AI Agents: How Zepto Scales Customer Support on Databricks and MLflow. Joint post with Zepto (Indian quick-commerce) whose customer support runs on a multi-agent AI system at 100,000+ tickets/day. Thesis: at scale the constraint shifts from model capability to system assurance (even 1% error = thousands of bad outcomes), so make evaluation the primary development primitive. The framework is a dual loop — a development loop (regression-test candidate versions against a golden dataset) + a production loop (monitor/score live traces) — joined by a quality gate that promotes only a version meeting all pillar thresholds and beating the production baseline (patterns/snapshot-replay-agent-evaluation). Seven phases on MLflow 3: (0) tracing — every invocation emits OTel spans centralized via Unity Catalog into Delta (patterns/telemetry-to-lakehouse); (1) evaluation pillars (customer experience / operational efficiency / risk & compliance / financial impact) with numeric gates; (2) golden dataset grown 500→2,000→5,247 examples, dev–prod accuracy gap 8→0.4 pts over six months (incl. security-team adversarial concepts/prompt-injection cases); (3) automated prompt optimization (reflect with a strong model, score with a cheap one — patterns/prompt-optimizer-flywheel, patterns/cheap-approximator-with-expensive-fallback); (4) scorers / "AI jury" — LLM judges + rules calibrated to human labels at 80–90% agreement, multi-judge for high-stakes (patterns/human-calibrated-llm-labeling, concepts/llm-as-judge); (5) model optionality over Model Serving / Foundation Model API (patterns/ai-gateway-provider-abstraction); (6) auto-regression (change → eval vs golden dataset + baseline → gate → auto-promote/reject); (7) production loop using risk-stratified sampling (new pattern patterns/stratified-evaluation-sampling) — effective 18–20% sample (~14,400 traces/day), 45–60% edge-case capture, 4–6 min detection, 86% lower review cost / 9× better edge-case detection vs uniform. Architecture is a decomposable orchestrator/router over vertical intent-specialists (WIMO/Missing/Expiry/Returns/Quality/Unable-to-Pay/General) + horizontal oversight agents (image dedup, item-match & manipulation detection) (patterns/specialized-agent-decomposition) so metrics compute per-agent. Four production stories: stale-cache ETA loop caught in 5 min → new feature line; WIMO cancellation regression caught pre-prod (92.1→87.4→94.2% intent, zero rollbacks); multimodal produce-scoring calibrated with Cohen's Kappa as reliability ceiling (concepts/multi-modal-attribute-extraction); refund-abuse image jury of 3 vision models (patterns/multimodal-content-understanding). Big lesson: never optimize a single metric (+5 intent alone → +133% latency, −0.4 CSAT; composite → +3 intent, +17% latency, +0.2 CSAT). Outcomes (self-reported): 80%+ AI-managed, 65% cost cut, payback <1 month, +20% CSAT, 3× dev speed, 4× resolution, ~155 eng-hours/month saved. New page: patterns/stratified-evaluation-sampling (≥2 sources incl. Figma Response Sampling as the uniform baseline). Extended systems/mlflow + 6 patterns + 6 concepts. Tier-3 Databricks + Zepto but on-scope: real multi-agent orchestration, trace-to-lakehouse observability at scale, and a stale-cache production incident. Caveat: all outcome numbers self-reported/unaudited; orchestrator internals and model roster undisclosed.
-
2026-09-08 — sources/2026-09-08-databricks-build-durable-agents-with-temporal-and-lakebase — Build durable agents with Temporal and Lakebase. Reference implementation of a durable long-running AI agent (personal-loan underwriting: gather evidence → apply governed policy → recommend → wait days for a human) pairing Temporal for durable execution with Lakebase Postgres for queryable operational state. Core design = a two-store split with a projection contract: Temporal Event History drives replay (Workflow / Activity / Signal), while Lakebase holds the application-facing view (run status, transcript, tool calls, review records, metrics) as an async-projected read model. The two do not share a transaction — Lakebase writes are Temporal Activities under at-least-once execution, so idempotency is enforced via deterministic IDs + Postgres PK/unique constraints + guarded upserts (compare-and-set on relational rows; retry against a terminal row affects zero rows). Governed underwriting policy flows from Unity Catalog via a continuous synced table (concepts/policy-as-data, no Worker redeploy); Change Data Feed (
REPLICA IDENTITY FULL, ~15-s batches) is the CDC return path to UC Delta history tables (lb_<table>_history) for audit.workflow.wait_conditionholds an open Workflow across a days-long human review with no Worker occupied — a long-running non-blocking wait step; the Workflow independently rejects stale/duplicate review Signals. OAuth M2M client refreshes its SQLAlchemy pool before the 1-hour DB credential expires (concepts/oauth-token-lifecycle). Per-op Retry Policies: model 4/3-min, tool 3/60-s, Lakebase 5/15-s. No new pages — the article maps entirely into existing canonical concepts/patterns/systems; extended systems/temporal, systems/lakebase, and 9 concept/pattern pages. Caveats: fixtures only (no lending-model/regulatory validation); crash test ran with Lakebase disabled; CDF still needs manual enablement; guarded-write wrapper doesn't yet classify zero-row results as failures. -
2026-09-07 — sources/2026-09-07-databricks-the-40-year-old-database-rule-agents-just-broke-how-ltap-unifies-oltp-and-olap-workloads — The 40-year-old database rule agents just broke: how LTAP unifies OLTP and OLAP workloads. Interview with Jonathan Katz (Postgres contributor, Databricks Senior Staff PM) supplying the motivation for LTAP (the companion to Matei Zaharia's 2026-06-30 storage-mechanics post). Thesis: the OLTP/OLAP split is 40 years old and rooted in storage physics (rows = one fast answer; columns = fast scan/aggregate), and AI agents are the first workload that can't tolerate the pipeline-and-copy delay. Worked failure mode: fraud detection clears in hundreds of ms, so a batch copy minutes old is useless — but a full-purchase-history scan against the live OLTP engine is an expensive analytical query that degrades every concurrent transaction (concepts/noisy-neighbor); a fleet of agents overwhelms it without guardrails. LTAP unifies at the storage layer, not the engine: Lakebase's PageServer transcodes Postgres rows into columnar Parquet "without changing a single bit," running a hot row tier + cool columnar tier (concepts/storage-media-tiering) so Spark/SQL read live operational data with no CDC and no OLTP load. Why it beats HTAP: "storage is the cheap part… compute is the expensive part" → serverless operational compute + serverless analytical compute scaled independently (concepts/compute-storage-separation); all under one Unity Catalog governance boundary (concepts/governed-agent-data-access); collapses the bronze/silver ingest tax (concepts/medallion-architecture). Slogan: "bring the engine to the data." New concept page: LTAP (canonical, now ≥2 sources). Extended: systems/lakebase, concepts/oltp-vs-olap, concepts/compute-storage-separation, concepts/columnar-storage-format, concepts/storage-media-tiering. Tier-3 Databricks interview but on-scope: real database/storage-layer architecture. Caveat: thought-leadership format, quantitative claims live in the companion Zaharia post; LTAP "rolling out," maturity asserted not benchmarked.
-
2026-09-04 — sources/2026-09-04-databricks-achieving-extreme-efficiency-through-specialized-gpu-kernel-generation — Achieving Extreme Efficiency through Specialized GPU Kernel Generation. Databricks' Proteus agentic harness generates GPU inference kernels specialized per runtime shape (static model params × dynamic per-request token count) rather than using generic kernels — reaching 1.8–5.2× over vLLM on Qwen 3.5 122B (Triton backend, NVIDIA B200). The durable systems lessons are about the harness, not the generator: (1) generation is cheap, validation is the bottleneck ("the system moves as fast as it can trust a kernel, not as fast as it can write one"); (2) agents reward-hack the harness (leftover-compile reuse, CUDA-graph replay mismatch, visible-test overfit) → defend with controlled-reference differential timing (CUDA-event + wall-clock + CUPTI, cleared state, re-timed winners), hidden test sets, and a >100× physically-impossible-speedup guard; (3) the knowledge layer can become the whole token cost → keep only actionable, scoped lessons retrieved by hierarchical tag + hybrid search, push distillation to background jobs; (4) shape-specialized search (case study: Gated DeltaNet packed decode, 0.025→0.018 ms, 1.5–1.6×); (5) intended redesign = "agent writes, loop validates". New system Proteus; new concepts reward-hacking-in-agentic-search, validation-as-the-bottleneck, actionable-scoped-lesson, program-search; new patterns controlled-reference-differential-timing, hidden-test-set-anti-overfit, physically-impossible-speedup-guard, shape-specialized-kernel-search, background-lesson-distillation, agent-writes-loop-validates-split. Extended systems/triton-lang, systems/vllm, systems/qwen, concepts/context-engineering, density-frequency-knowledge-partition. Tier-3 Databricks but on-scope: real serving-infra + agent-harness architecture. Caveat: speedups are per-kernel, not end-to-end; the "agent writes, loop validates" split is aspirational.
- 2026-09-03 — sources/2026-09-03-databricks-governance-beyond-security-knowledge-context-ontology-on-the-lakehouse — Governance beyond security: knowledge, context & ontology on the lakehouse. Databricks Data Empowerment Program (DEP) post (Tier-3, borderline-include: ~50%+ architecture) reframing AI data governance as knowledge, context, and ontology, not just security. Durable ideas: (1) catalog-as-agent-runtime — governance artifacts (tags, comments, certified flags, lineage, glossary terms, de-id policies, contracts, model cards) live in Unity Catalog as machine-readable metadata that agents execute against ("the catalog isn't where you document governance; it's the runtime the agents execute against"); (2) read → act → write-back loop — five build agents (ETL-gen, testing, curation-validation, de-id, deploy) read instructions and write evidence (test outcomes, quality scores, lineage, CDC) back, De-ID+Testing sequenced first, human shifts from authoring metadata to approving it; (3) ABAC at the data layer spanning SQL + vector search — "if a user cannot query a row in SQL, no agent can retrieve it via vector search or embeddings," agents act with the querying user's entitlements (not a service account); (4) fail-closed suppression of anything the catalog doesn't describe; (5) De-ID from security-policy metadata (Discover→Curate→Execute; HIPAA Safe Harbor, referentially-intact synthetic data; non-prod boundary isolation so PHI never leaves the governed boundary); (6) one analytic agent bound to one certified product — "single largest accuracy lever," wrong answers route to the named Data Product Owner, certified metrics computed directly (enforced, not documented); (7) AI Certification as a queryable scorecard with continuous expiration (Governance/Quality/Semantics auto-scored, Ownership steward-signed) rolled into an AI-readiness score; (8) governance-as-AI-foundation + data-readiness-vs-model-spend — "govern the data well enough, and AI can run on cheaper models with more trust … the intelligence lives in the catalog, not the token bill." New concepts: catalog-as-agent-runtime, governance-as-ai-foundation, data-readiness-vs-model-spend, unified-permission-model-over-data-and-embeddings. New patterns: catalog-metadata-as-agent-instruction-set, de-identification-from-security-policy-metadata, agent-bound-to-single-certified-data-product. Extended: attribute-based-access-control, fail-closed-on-sensitive-data, run-with-end-user-credentials, data-centric-ai-governance, data-products-as-ai-context. Tier-3 DEP/advocacy post, healthcare-framed; the "cheaper model, identical accuracy" claim is asserted not benchmarked; distinct from the same-day Building High-Quality Data Products post (that = lifecycle+contract; this = the agentic runtime + model-economics).
- 2026-09-03 — sources/2026-09-03-databricks-building-high-quality-and-trusted-data-products-with-databricks — Building High-Quality and Trusted Data Products. A guidance/advocacy post (Tier-3, borderline-include for its portable design content) arguing enterprises should treat data as a product — a first-class owned asset — whether or not they adopt data mesh. The durable ideas: (1) the five characteristics every data product must meet (quality & observability, semantic consistency, privacy, security, discoverability); (2) a required data product owner accountable across the lifecycle; (3) the seven-phase lifecycle (inception → design → creation → publish → operate/govern → consume → retirement) with the sharp publish ≠ create distinction (publishing adds deploy, catalog registration, permissioning, and change-versioning as its own governed step); (4) the data contract as the federated-governance mechanism — a producer↔consumer agreement carrying description, schema/formats, usage policies, quality, security, data SLAs, and responsibilities incl. the change process; (5) certification as a governed-publication workflow where a steward reviews the proposed contract, CI/CD then deploys production pipelines, and the catalog records status as tags + markdown (metadata, not a runtime gate); (6) a cross-functional governance team as Center of Excellence standardising contract framing. Databricks mapping: Delta Live Tables expectations validate quality at each medallion layer, Auto Loader lands Bronze, Unity Catalog governs/lineages/tags/discovers, Lakehouse Monitoring watches the contract, Delta Sharing / Marketplace distribute. Customer numbers (marketing-sourced): Rivian 25k vehicles TB/day + 30–50% runtime; Walgreens 825M scripts/yr, 40k events/sec; T-Mobile mesh via UC + Delta Sharing. New concepts: data-as-a-product, data-product-certification. New pattern: data-contract. Extended: concepts/data-mesh, concepts/medallion-architecture, concepts/data-products-as-ai-context, concepts/certification-as-metadata-not-enforcement. Tier-3 guidance post; ideas predate it (DJ Patil's Data Jujitsu + mesh literature cited), vendor feature-mapping and customer stats not independently verified.
- 2026-09-01 — sources/2026-09-01-databricks-how-we-eliminated-1-million-a-year-of-wasted-ai-agent-spend — How we eliminated $1M/year of wasted AI agent spend in one hour. Databricks agents call MCP tool servers (Jira, GDrive/Docs); when a tool misbehaves the agent doesn't fail loudly — it retries and works around it, so the task completes while tokens + wait time burn silently (silent-tool-failure-cost). Because Unity Gateway sits inline it already emits one OpenTelemetry trace per MCP call (tool name, args, error, tokens, latency, session ID) into a single unified trace table; pointing Genie One at it answered "which errors recur / how many turns to recover / what does each cost" in plain English (trace-driven-tool-failure-diagnosis). Result: seven tool bugs = ~$499K/yr tokens + ~12,000 eng-hours/yr (~$1.2M/yr), found+fixed in ~1 hour. Recovery cost tracks error-message quality almost perfectly (self-documenting ~14% repeat / 4.6 turns; cryptic traceback 30.5% / 12.1; misleading 50% / 13.1 — worse than cryptic). Key reframe: the model usually wasn't calling tools wrong — signatures are kept deliberately loose (to save per-call token cost), the model fills gaps with a reasonable guess (a JSON array for "list of fields" → the
.split()bug, 535 fails/day, $87K/yr), and the tool crashed on every shape but one. Design principle: build tools that adapt to how LLMs naturally call them (coerce, default, absorb). New concepts: silent-tool-failure-cost, error-message-quality-drives-agent-recovery-cost, under-specified-tool-signature. New patterns: trace-driven-tool-failure-diagnosis, tools-adapt-to-how-llms-call-them. Extended: systems/unity-ai-gateway, systems/databricks-genie, systems/opentelemetry, systems/model-context-protocol, concepts/token-overhead. First-party engineering blog; $/hour figures are Databricks' own estimates extrapolated from a single 24-hour trace window, not an audited study. - 2026-09-01 — sources/2026-09-01-databricks-collaboration-makes-us-all-stronger — Collaboration makes us all stronger — a coordinated-disclosure story around a memory-safety bug in the
address_standardizersub-extension of PostGIS, shipped to tenants on Lakebase and Neon. The bug: a caller-controlled grammar-rule value indexes a fixed-size internal array without a bounds check → out-of-bounds access, reachable by a normal tenant role (no privilege needed). Two architecture points carry the post: (1) Databricks runs Lakebase/Neon on a microVM architecture whose compute-instance boundary meant the exploit caused no cross-customer impact on Databricks (vs cross-customer exposure the researcher saw elsewhere) — blast-radius containment; (2) their extension build system applies an arbitrary patch set on top of any upstream extension before compile (backport a fix or disable a risky code path), a downstream patch lever that protects customers independent of the upstream release timeline. Response arc: Neon production alarms detected the PoC testing → a security engineer proactively reached out → hardened fix deployed to Lakebase + Neon tenants (first pass incomplete → +1 week, no customer action) → root-cause fix driven upstream into PostGIS; the upstream fix landed as a CVE-less "minor" memory-leak cleanup and was itself incomplete until the researcher submitted the rest. Core thesis: own your exposure, not just your code — a bug in an OSS component you ship to untrusted input is your bug. New concept: own-your-exposure-not-just-your-code. New pattern: downstream-patch-lever-for-shipped-oss. Extended: systems/postgis, systems/lakebase, systems/neon, concepts/memory-safety, concepts/microvm-session-isolation, patterns/upstream-the-fix (ninth instance), patterns/postgres-extension-over-fork (shared-attack-surface flip side). First-party security blog; exploitation detail deferred to the researcher's deep-dive, "no cross-customer impact" is a first-party claim, competing-platform exposure referenced not characterized. - 2026-08-28 — sources/2026-08-28-databricks-fast-fault-tolerant-pytorch-training-on-ai-runtime — Fast, fault-tolerant PyTorch training on AI Runtime. Argues large-scale training efficiency reduces to one metric — goodput (fraction of GPU time on productive compute) — and that GPU failures are the expected case (~1% annualized per-GPU rate → 19% failure chance for a 256-GPU/30-day job, 57% at 1,024 GPUs; the 608×H100 Delta supercomputer fails every ~1.9 h). Four compounding levers: (1) use PyTorch Distributed Checkpoint (DCP) — every rank writes its shard in parallel (save time ~1/N), a
.metadatafile records global layout so checkpoints reshard onto a different GPU count, and it's worth it even for DDP; (2) async saves (async_save: fast staging copy + background upload) make frequency nearly free — measured 1.8× (DDP 2.8B) / 58× (FSDP 20B) vstorch.saveon 32×H100; (3) automatic recovery to the latest complete checkpoint (.metadatawritten only after all shards land); (4) overlap dataloading with compute and cache remote data to local NVMe —UCVolumeDatasetover UC Volumes gives 20–50% wall-clock reduction (image example: GPU util 12.6%→53.3%, epoch-2 throughput 371→6,590 img/s), withfetch_secondslogged to MLflow. Fifth, subtler point: checkpoint the data-pipeline position + RNG state or restarts silently corrupt the training distribution. Checkpoint-frequency arithmetic (Llama 3 ~8.6 interruptions/day): 2 h interval → 64% goodput, 30 min → 91%. New concepts: goodput, distributed-checkpoint, checkpoint-frequency, data-pipeline-checkpointing. New system: pytorch-distributed-checkpoint. New patterns: asynchronous-checkpoint-staging, overlapped-dataloading, local-nvme-cache-for-remote-training-data, resume-to-latest-complete-checkpoint. Extended: systems/pytorch, systems/mlflow, systems/unity-catalog-volumes, concepts/gpu-training-failure-modes, concepts/annual-failure-rate, concepts/gpu-stall-from-storage. Mechanisms-and-tradeoffs post; benchmarks vendor-reported (async numbers exclude thetorch.savebaseline's network-storage time),UCVolumeDatasetpartitioning/eviction params undisclosed. - 2026-08-27 — sources/2026-08-27-databricks-building-for-the-ai-era-lakebase-streaming-and-lakehouse-innovations-vldb-2026 — VLDB 2026 preview roundup framed by Reynold Xin's "three golden ages of database engineering" keynote thesis (AI agents drive the third; answers are LTAP + Lakebase). Four accepted papers + a demo + a sponsor talk: Lakebase (serverless Postgres over open lake storage, sub-second cold starts, copy-on-write branching for "millions of short-lived, deeply branched databases"; Stas Kelvich); Spark Structured Streaming — a decade of evolution powering millions of weekly jobs, micro-batch pipelining up to 3× throughput, new stateful APIs, fine-grained access control (Siying Dong); AutoLiquid — autonomic clustering-key selection via
CLUSTER BY AUTOwith heuristics + shadow verification, beating customer-selected keys on >95% of workloads (Yunjia Zhang); Ultron — history-based query optimization exploiting workload repetitiveness, 25% median join-latency improvement (Eric Liang); Enzyme IVM demo (Yuhong Chen); and the sponsor talk introducing LakehouseRT (new Reyden engine) for real-time analytics over open storage — Lakebase + LakehouseRT positioned as "the first true LTAP system" (Ippokratis Pandis). New systems: autoliquid, ultron, lakehousert. New concept: history-based-query-optimization. New pattern: shadow-verification-for-autonomic-optimization. Extended: systems/lakebase, systems/spark-streaming, systems/liquid-clustering, systems/enzyme-ivm, concepts/ltap. Announcement/preview depth — paper internals (Ultron history store, Reyden execution model) named but not detailed; numbers vendor-reported. - 2026-08-24 — sources/2026-08-24-databricks-how-databricks-uses-ai-to-accelerate-incident-investigation — AI SRE, Databricks' AI-powered incident-investigation agent for a fleet of 100s of microservices on 1500+ Kubernetes clusters across 70+ regions and three clouds. Design thesis from dozens of on-call interviews: context assembly consumes 60–80% of investigation time, not root-cause insight — so AI SRE is built as a context-assembly engine, not a chatbot on top of observability. Automatic triage launches three tracks in parallel the instant an incident fires — platform health checks (cloud/network/upstream deps/Auth, eliminating a whole class of red herrings), service-level analysis (logs/metrics/traces/deploys/config, anomalies vs baseline), and team-owned agentic runbook execution — and correlates them into one diagnostic summary before the engineer opens their laptop; plus an interactive natural-language debugging mode. Built as a four-tier layered platform (Primitives → API Layer → Core Engine → Application Layer) so first-party bots and team runbooks share one substrate. Three reliability principles: deterministic structured checks before open-ended LLM reasoning, every conclusion links to auditable evidence, and graceful degradation (honest partial answer over hallucinated diagnosis). Notable systems lesson: guardrails matter more for agents than people — agents hit endpoints in bursts and never self-throttle, so the API layer was redesigned to protect shared observability infra. Augment-not-replace; investigation-only today. Impact: 150+ teams, 250+ WAU, 2,000+ investigations/day. New system: databricks-ai-sre. New patterns: layered-debugging-platform, agentic-runbook-as-team-owned-skill. New concepts: structured-checks-before-llm-reasoning, traceable-evidence-for-agent-trust, graceful-agent-degradation, agent-api-guardrails. Extended: concepts/on-call-automation, concepts/agent-as-first-pass-investigator, patterns/evidence-fan-out-then-synthesize, systems/instacart-blueberry. First-party engineering blog; self-reported impact, no controlled MTTR study; specific LLM(s) and orchestration internals undisclosed.
- 2026-08-20 — sources/2026-08-20-databricks-busting-sql-migration-myths-how-new-sql-features-make-lift-and-shift-to-lakehouse-easier — Lift-and-shift of a nightly Oracle stored procedure onto the Lakehouse by translation, not rewrite. Mostly a SQL-scripting syntax-mapping how-to (legacy
BEGIN ... EXCEPTION→DECLARE EXIT HANDLER FOR SQLEXCEPTION; native cursorsOPEN/FETCH/CLOSEsince Runtime 18.1;CREATE TEMP TABLE; full control-flow toolkit;MERGEas-is). Ingested (borderline-include, Tier 3) for its one architectural nugget: the transaction concurrency model.BEGIN ATOMIC ... ENDgives all-or-nothing commit/rollback with row-level conflict detection — "Concurrent batches writing to the same table only conflict if they touch the same rows. For instance, Oracle and Snowflake both use table-level locking, which forces serial execution." Gated on the DeltacatalogManagedtable feature (ALTER TABLE ... SET TBLPROPERTIES('delta.feature.catalogManaged'='supported'));BEGIN ATOMICmust sit at top level. Post-migration payoff is Unity Catalog registration (SQL SECURITY {INVOKER|DEFINER}, access controls, column-level lineage, discoverability). Claimed 50–75% migration-timeline reduction (vendor claim). New concept: table-level-vs-row-level-locking. Extended: row-level-concurrency, concepts/optimistic-locking, concepts/open-table-format, systems/delta-lake, systems/unity-catalog, systems/databricks-sql-warehouses. Vendor competitor-comparison claim (Oracle/Snowflake table-level locking) uncited; conflict-detection implementation undisclosed. - 2026-08-17 — sources/2026-08-17-databricks-how-databricks-feature-store-serves-features-with-sub-second-freshness — Databricks Feature Store's sub-second streaming feature-serving architecture. Author a feature once → the same definition drives offline batch (baselines/training) and online streaming (fresh signals). Online path: Kafka → Spark Real-Time Mode rolling aggregations → Lakebase via a streaming JDBC sink → Model Serving auto-lookup at inference, hitting 200 ms end-to-end p99. Key internals: rolling windows (per-event, ms-resolution, always fresh) vs tumbling/sliding; RTM runs stages concurrently with per-row aggregate update + per-row expiry on a local RocksDB state store (state exceeds cluster memory); checkpoint cost amortized over longer intervals, exactly-once with ≤5-min Kafka replay; serverless on SDP with pre-provisioned near-zero-interruption restarts. Lakebase minimizes streaming write amplification vs stock Postgres by writing compact change records (safekeeper-quorum durability) instead of full 8 KB page writes on hot rows; online store scales to 10s-of-thousands reads/sec at 10s-of-ms; Model Serving 100K+ QPS CPU. Training/skew solved via offline Kafka copy + point-in-time joins, also used to backfill online features. Features are first-class Unity Catalog objects; MLflow records feature dependencies. New concept: rolling-window-aggregation. New patterns: author-once-serve-batch-and-streaming, streaming-jdbc-sink-to-online-store, backfill-online-features-from-offline-copy. Extended: concepts/feature-store, concepts/micro-batching, concepts/postgres-full-page-write, systems/apache-spark, systems/spark-streaming, systems/rocksdb, systems/lakebase, systems/databricks-model-serving, systems/mlflow, systems/unity-catalog. First-party product blog; 200 ms/QPS numbers not independently benchmarked.
- 2026-08-13 — sources/2026-08-13-databricks-smart-routing-in-unity-ai-gateway — Architecture of Smart Routing in Unity AI Gateway (Beta), the model-router the 2026-08-07 cost playbook only named. Central decision: task-aware routing over per-request routing — because coding-agent cost is "dominated by cache hit rate", the router commits a model+effort for the whole session (per-request routing shatters the prompt cache and is rejected). Internals: a cheap low-latency classifier labels the task (system area, code evidence, failure mode, fix localization, project type → task-type + language families), then the router triangulates from a medium default, escalating/delegating under a single policy (patterns/two-stage-evaluation). Routes on start-of-task info only. Runs across models and harnesses in Omnigent (patterns/specialized-agent-decomposition); all sub-agent launches go through the Smart Routing API for fresh-cache re-routing. Roadmap: route where scoping is free first (PR reviews, sub-agent launches, batch jobs), route interactive work after a few turns, switch at the context-compaction seam by pricing cache misses (à la Cognition Devin Fusion). Traces logged to Unity Catalog for an AI+human eval loop. Results: 35% (internal) / 56% (public-benchmark) savings; outperformed any single model at 65% of Opus 5's cost/task; matched Opus 5 at <50% cost. New concept: task-aware-vs-per-request-routing. New patterns: two-stage-classifier-then-router, route-at-context-compaction-seam. Extended: systems/unity-ai-gateway, systems/omnigent, patterns/classifier-based-smart-routing, patterns/task-level-routing-via-meta-harness, patterns/request-level-model-routing (corrected: Smart Routing is task-aware, not per-request), patterns/auto-escalation-on-quality-failure, concepts/cache-hit-rate. Tier-3 Beta launch, ingested as borderline-include (real serving-infra routing architecture > 20% of body); classifier model/accuracy/latency undisclosed.
- 2026-08-12 — sources/2026-08-12-databricks-how-amtrak-is-building-the-data-backbone-for-its-largest-transformation — Customer case study (Amtrak). Rail Intelligence, a unified data backbone on the Databricks lakehouse for Amtrak's largest physical transformation in 50+ years (186 mph NextGen Acela ×28, 83 Siemens Airo trainsets across 14 corridors, 21,000 track miles; each trainset 100+ sensors). Architecture shape: fan-in via Lakeflow Connect + real-time streaming (fleet IoT, wayside detectors, dispatch, geospatial, Sqills S3 Passenger reservations, OT) → Delta Lake bronze → medallion refinement → trusted data products governed by Unity Catalog (lineage, domain access control, data-quality contracts) → ML (anomaly detection, computer-vision defect pipeline, delay-probability) via MLflow + Model Serving → maturity endpoint of agentic/NL queries through Genie embedded in a Databricks Apps experience layer. Five intelligence products (Fleet Health / Safety / Operational Readiness / Reservations / Capital Prioritization; $5.5B/yr capital program). Strategic thesis: a compounding data platform — every trainset, capital project, and booking adds a signal that improves every model. New concepts: data-network-effect, predictive-maintenance. Extended: systems/delta-lake, systems/unity-catalog, systems/mlflow, systems/databricks-model-serving, systems/databricks-genie, systems/databricks-apps, systems/lakeflow-jobs, concepts/medallion-architecture. Case-study-depth (reference architecture + business numbers), no system internals.
- 2026-08-12 — sources/2026-08-12-databricks-network-configuration-delivery-to-tens-of-millions-of-serverless-vms — Re-architecting network-config delivery for tens of millions of serverless VMs/day (billions of config requests/day). The old design synchronously fanned out to all upstream services on the cluster-launch critical path — ~5,000 ms p99, 99.8% from compound availability. New design: upstreams emit identifier-only change events to a message queue → background pipeline recomputes and materializes a per-workspace snapshot → launch path does a single read (control/data-plane split), with a low-frequency reconciler as backstop. Results: 125 ms p99 (−97.5%), 99.99%, −86% upstream calls; upstream outages degrade to static config. New system: databricks-network-config-delivery. New concepts: snapshot-precomputation-serving-path, reconciler-as-safety-net. New patterns: event-driven-snapshot-precomputation, partition-colocated-precomputation. Extended: concepts/event-driven-architecture, concepts/static-stability.
- 2026-08-11 — sources/2026-08-11-databricks-taking-auto-cdc-to-the-next-level-solving-the-hardest-real-world-use-cases — Follow-up extending AUTO CDC beyond SCD 1/2/snapshot to three hard cases: bitemporal history (
STORED AS BITEMPORAL, four system columns__START_AT/__END_AT+__SYSTEM_START_AT/__SYSTEM_END_AT, dualSEQUENCE BY/SYSTEM SEQUENCE BY, Beta) for two-clock point-in-time reconstruction under SEC 17a-4 / FINRA (>$2B fines since 2021); Partial Updates GA where NULL means "do not update" (IGNORE NULL UPDATES ON/* EXCEPT/COLUMNS TO UPDATE); and reproducible ML that survivesVACUUMbecause history is stored as data rows, not reclaimable Delta file versions (reproducibility-survives-vacuum). Plus open-sourcing the AUTO CDC Type 1 Python API in Apache Spark 4.2 (auxiliary state table for out-of-order events, idempotent-retry convergence, runs on both Delta and Iceberg). New concepts: bitemporal-modeling, point-in-time-reconstruction, partial-update-cdc, reproducibility-survives-vacuum. New patterns: null-means-do-not-update, history-as-data-over-file-versioning. Extended: systems/databricks-autocdc, systems/delta-lake, systems/apache-spark, concepts/change-data-capture, concepts/time-travel-data-state. - 2026-08-11 — sources/2026-08-11-databricks-open-sourcing-metals-v2 — Open-sourcing Metals v2, Databricks' Java + Scala language server for a 26M-line Bazel monorepo (2.9M symbols, 142k+ files). New north-star metric: time-to-initial-intelligence (TTII) (p50 8.7s, p90 36.7s). Three reworked layers: (1) build-free repo index — the mbt index, a Git-blob-OID content-addressed index with per-document bloom filters, removing BSP from the startup path; (2) compiler-backed pipelines — one Scala presentation compiler holding 24M lines via outline mode + fallback/precise modes, and Java on
javac+ Google Turbine (~1M lines/sec); (3) metadata-first BSP querying 285k Bazel targets off the hot path. Adoption: 92% weekly IDE users on Cursor vs 12% IntelliJ; Stripe early adopter. New systems: metals-v2, metals-mbt-index, google-turbine, build-server-protocol. New concepts: time-to-initial-intelligence, language-server-protocol, content-addressed-index, presentation-compiler. New patterns: build-free-repo-index, outline-mode-type-checking, metadata-first-build-integration, content-hash-incremental-index, fallback-then-precise-compiler-mode. Extended: systems/bazel, concepts/monorepo, concepts/incremental-indexing, concepts/bloom-filter. - 2026-08-10 — sources/2026-08-10-databricks-how-to-ground-genie-agents-in-both-structured-data-and-documents-without-losing-governance — How Genie Agents ground on both structured tables and documents while keeping Unity Catalog — not the LLM — as the security perimeter. Core principle: Genie Agents run with the end user's credentials, so every answer is filtered at the data layer before leaving the Lakehouse; prompt-layer filtering "makes the LLM your security perimeter, a dangerous bet." Four steps: (0) identity sync via AIM + JIT; (1) four structured-data access layers — object privileges, ABAC, row filters, column masks; (2) documents governed via UC Volumes (attached volume = required source,
READ VOLUMEgates agent use, smallest securable unit → one-audience-per-volume); (3) test by impersonation, not inspection. New concepts: run-with-end-user-credentials, identity-sync-as-governance-foundation, test-governance-by-impersonation. New pattern: volume-as-agent-knowledge-source-with-required-access. Extended: systems/databricks-genie, systems/unity-catalog, systems/unity-catalog-abac, systems/unity-catalog-volumes, concepts/attribute-based-access-control, concepts/session-identity-evaluation, concepts/governed-agent-data-access, concepts/governance-travels-with-resources, patterns/tag-driven-attribute-based-access-control. - 2026-08-07 — sources/2026-08-07-databricks-managing-ai-coding-costs-at-scale — Industry cost-management playbook for AI coding tools (Databricks + Stripe, Coinbase, Uber, Ramp). Five levers: chase the efficiency frontier (best price for a given intelligence level, advances faster than the intelligence frontier), preserve model flexibility via a meta-harness (Omnigent), dynamic routing (request-level proxy / task-level meta-harness / escalation-delegation), progressive spend friction (visibility → gates → downshift → suspend) over hard budgets, and cut token overhead (compaction, cache tuning). All five presuppose the AI Gateway (Unity AI Gateway Smart Router: >30% avg task-cost reduction at matched quality; harness/cache tuning: ~50% token reduction). New concepts: efficiency-frontier, token-overhead, ai-gateway-design-pattern. New patterns: request-level-model-routing, task-level-routing-via-meta-harness, progressive-spend-friction, prompt-cache-tuning-for-cost. New system: meta-harness. Extended: systems/unity-ai-gateway, systems/omnigent, patterns/auto-escalation-on-quality-failure.
- 2026-07-28 — sources/2026-07-28-databricks-coding-agent-spend-unity-ai-gateway-budgets — How Databricks manages internal coding agent spend with Unity AI Gateway Budgets. Hard enforcement now live: dual-budget architecture (daily runaway guard + monthly cap), self-serve acknowledgment workflow, tiered overrides via group membership. 500–1,000 engineers hit old single-limit monthly; new model eliminates interrupt queue. New concepts: dual-budget-spend-control. New patterns: dual-limit-daily-plus-monthly, self-serve-acknowledgment-workflow, tiered-override-via-group-membership. Extended: systems/unity-ai-gateway-budgets, systems/unity-ai-gateway.
- 2026-07-23 — sources/2026-07-23-databricks-self-serve-infrastructure-vending-machine — Databricks' Field Engineering Vending Machine (FEVM): self-service infrastructure provisioning platform built as a Databricks App with Terraform + Lakebase + MCP. Use-case-based provisioning, TTL-based lifecycle, agent-first API. 5,000+ users, 2,600+ active deployments across 3 clouds, ~1,200 requests/day burst. New: systems/databricks-fevm, concepts/self-service-infrastructure, concepts/resource-lifecycle-management, patterns/self-service-infrastructure-vending-machine, patterns/use-case-based-provisioning, patterns/ttl-based-resource-lifecycle, patterns/agent-first-infrastructure-api. Extended: systems/databricks-apps.
- 2026-07-23 — sources/2026-07-23-databricks-intent-based-authorization-omnigent — Introduces intent-based authorization in Omnigent: binds agent sessions to a declared purpose, evaluates every tool call against it (ALLOW/ASK/DENY), blocks indirect prompt injection attacks that exploit identity-based authorization's purpose-blindness. Deny-wins policy composition, session-scoped immutable intent, human-gated activation. New: systems/omnigent, concepts/intent-based-authorization, concepts/confused-deputy-problem, patterns/intent-based-authorization, patterns/deny-wins-policy-composition, patterns/session-scoped-intent-binding, patterns/human-gated-tool-invocation.
- 2026-07-22 — sources/2026-07-22-databricks-simplify-ai-agent-orchestration-with-lakebase-postgres — CLA/Databricks Forward Deployed Engineering: Lakebase Postgres as sole orchestration backbone for production document-processing agent system. Four Postgres-native patterns (FOR UPDATE SKIP LOCKED priority dequeue, lease-based crash recovery, rate-limit-aware throttling, idempotent webhook callbacks) replace Kafka/Redis/Airflow/Temporal. LISTEN/NOTIFY + SSE for real-time dashboard. New patterns: lease-based-task-recovery, listen-notify-push-with-polling-fallback, rate-limit-aware-dequeue-throttling. Extended: select-for-update-skip-locked, postgres-queue-on-same-database, lakebase, idempotent-operations.
- 2026-07-21 — sources/2026-07-21-databricks-why-rd-data-belongs-in-the-lakehouse-and-why-agents-need-it — cellcentric (Daimler Truck × Volvo) case study: four-year lakehouse investment becomes governed AI context layer. Data Hub with dual interface (marketplace UI + MCP server), 27 data products, context-coverage-as-quality-metric, OAuth 2.0 on-behalf-of identity delegation for agents, hybrid MCP (Databricks-managed + custom), continuous evaluation framework (MLflow tracing + multi-scorer eval). New concepts: data-products-as-ai-context, context-coverage-as-quality-metric, agent-governance-via-identity. New patterns: data-product-as-agent-context-layer, dual-interface-human-and-agent, continuous-agent-evaluation-framework. Extended: on-behalf-of-agent-authorization, unity-catalog, data-lakehouse, data-mesh.
-
2026-07-16 — sources/2026-07-16-databricks-what-happens-in-the-milliseconds-after-you-tap-pay (Tier-3; passes scope: real-time fraud detection architecture with sub-50ms end-to-end latency. Demonstrates route optimization as data-plane shortcut, OAuth-authenticated connection pooling with double-check-locking pool rebuild, and latency waterfall observability. Production benchmark: 27 ms p50, 37 ms p95.)
-
2026-07-13 — sources/2026-07-13-databricks-ultra-fast-anomaly-detection-spark-rtm (Tier-3; passes scope: streaming architecture with Spark Real-Time Mode demonstrating sub-millisecond P95, 1 ms P99, 69,713 rows/sec sustained on ~23M records. Names three RTM architectural innovations (continuous data flow, pipeline scheduling, streaming shuffle). Introduces guardrail-stream-pipeline and single-trigger-latency-class-switch. Production validation by Coinbase, DraftKings, MakeMyTrip.)
-
2026-07-06 — sources/2026-07-06-databricks-scaling-security-alert-triage (Tier-3; passes scope: production agent fleet architecture for security alert triage — 17 source-specific agents on Structured Streaming, deterministic filtering handles 30–95% of volume, three-layer cost controls, MLflow tracing + analyst ground-truth evaluation. Key numbers: 18K+ alerts triaged, 3.2% escalation rate, 10.5s median, 6,500+ analyst hours saved in 30 days.)
-
2026-07-01 — sources/2026-07-01-databricks-gpu-reliability (Tier-3; passes scope: GPU fleet reliability infrastructure — multi-stage health check architecture (systems/gpu-monitor), NCCL timeout mechanics (nccl-ib-timeout), cumulative-downtime monitoring, node quarantine-and-retest pattern, inter-node fabric bandwidth validation at multiple payload sizes. Introduces systems/gpu-monitor, systems/nccl, systems/dcgm, nccl-ib-timeout, cumulative-downtime-vs-flap-count, multi-stage-health-check, node-quarantine-and-retest.)
- 2026-06-30 — sources/2026-06-30-databricks-from-monolith-to-lakebase-to-ltap (Tier-3; passes scope: deep database storage architecture — WAL externalization to SafeKeeper via Paxos, PageServer with object-storage backing, LTAP row-to-columnar transcoding eliminating CDC, MVCC preservation in columnar format, LSN-consistent analytical reads. Introduces ltap and storage-layer-unification-over-engine-unification.)
- 2026-06-12 — sources/2026-06-12-databricks-enabling-evolutionary-database-development-database-branchin-part3 (Tier-3; passes scope: team-scale database branching architecture — tier topology as long-running branches, permission model with policy-enforced governance, DBA-to-platform-engineer evolution, SCM state machine with blocking gates for agent governance, TDD layer with per-role agents. Introduces systems/lakebase-app-dev-kit.)
- 2026-06-11 — sources/2026-06-11-databricks-ingesting-the-milky-way-petabyte-scale-with-zerobus-ingest (Tier-3; passes scope: deep streaming architecture with dynamic partitioning internals, zero-copy protobuf decoder (Zeroparser, ~1 GB/s/core, Rust, OSS), latency-optimized WAL, petabyte-scale benchmark at 12 GB/s sustained / 1.04T rows / 2,048 streams)
- 2026-06-10 — sources/2026-06-10-databricks-ai-serving-platform-that-adapts-to-your-model (Tier-3; passes scope: model-serving infrastructure architecture with two-axis autoscaling internals, cold-start mitigation via warm pools, operational numbers at 300K+ QPS / 99.99% availability / 10→10K QPS in <60s, architectural comparison of request-based vs resource-based scaling)
- 2026-06-05 — sources/2026-06-05-databricks-enabling-evolutionary-database-development-database-branchin (Tier-3; passes scope: Part 2 of the Evolutionary Database Development series — CI/CD workflow architecture with GitHub Actions templates, new patterns: expand-and-contract, A/B variant prototyping, destructive testing; idempotent migration as hard requirement; DBA role as async PR reviewer)
- 2026-06-03 — sources/2026-06-03-databricks-apache-spark-real-time-mode-for-gaming
(Tier-3; passes scope: streaming architecture with stateful
processing,
transformWithStateoperator internals, production numbers at 4M sessions / 500K events-min / 432ms p99, architectural comparison vs Flink and custom actor systems) -
2026-06-01 — sources/2026-06-01-databricks-debunking-8-data-layout-myths-why-liquid-clustering-outperfo (Tier-3 Databricks Blog post; passes scope decisively despite consultative-listicle shape because the technical content is dense — transaction-log-based pruning mechanism, per-file min/max statistics powering both file skipping and metadata- only operations, OPTIMIZE engineering improvements with specific 12h→23min planning numbers, three production case studies with concrete numbers, structural critique of Z-Order rewrites, row-level vs file-level concurrency distinction, co-clustered join shuffle elimination). Debunking 8 data layout myths: why Liquid Clustering outperforms partitioning. Most architecturally dense Liquid Clustering disclosure on the wiki — each of eight defender-of-partitioning arguments paired with the verbatim reality, plus three production case studies at PB scale. Eight first-class wiki primitives canonicalised: concepts/horizontal-sharding (failure rate disclosed: "more than 75% of cases" on the Databricks customer base; three mistake classes — high cardinality, wrong column, over-fine granularity); concepts/open-table-format (load-bearing architectural fact: "directory-pruning does not exist on modern open table formats like Delta and Iceberg" — pruning is per-file via transaction-log statistics); metadata-only-operation (DELETE / COUNT / DISTINCT / GROUP BY computed from per-file min/max stats; ~90% faster metadata-only DELETEs, up to 27× aggregate speedups); row-level-concurrency ("two writers updating different rows no longer conflict, even if those rows live in the same file"; partition-as-write-boundary is a workaround for an older concurrency model); concepts/z-ordering (the older clustering technique with two structural problems — "poor clustering quality" + "unnecessary rewrites" — superseded by Liquid Clustering); multi-dimensional-clustering (clustering on
(date, hour, source, id)simultaneously, impossible under partitioning's cardinality limits); co-clustered-join (Private Preview shuffle elimination; ~51% faster, 87% less shuffle data on a real benchmark); low-cardinality-clustering-optimization (per-file=single-low-cardinality-value layout with high-cardinality nested sort; 35% lower clustering time, 22% faster query times benchmark). Four new patterns: clustering-keys-as-engine-input ("Liquid treats clustering keys as input that the engine uses to guide optimal file organization. Keys can be changed at any time" — the layout-as-implementation-detail thesis); incremental-clustering-on-write ("Liquid clusters incrementally, including at write time, so the layout stays optimal without unnecessary rewrites" — bounds maintenance cost to new-data volume rather than table size); in-place-partitioned-to-clustered-conversion (ALTER TABLE .. REPLACE PARTITIONED BY WITH CLUSTER BY— Private Preview, validated by Bolt's zero-downtime CDC migration); replace-using-and-replace-on-for-selective-overwrite (REPLACE USING / ON layout-agnostic, compute-agnostic — works on any layout and any compute, unlike partitioning-tied Dynamic Partition Overwrite). One new system page: systems/arctic-wolf-security-telemetry-table — production 3.8+ PB security telemetry table ingesting 1+ trillion events per day, 90-day queries 51s → 6.6s (7.7×) post-migration to Liquid Clustering on UC managed tables with Predictive Optimization. Three production case studies with operational numbers: Arctic Wolf (3.8+ PB; 1T+ events/day; 7.7× query speedup; file count 4M → 2M; data freshness hours → minutes); Bolt (TB-scale CDC; +138% write throughput; −21% avg / −63% max read time; zero downtime alongside live ingestion via in-place Liquid Conversion); Databricks-internal (1.1 PB → 0.8 PB / −27% storage; 5.9× wall-clock speedup; 86% bytes- read reduction across 16 representative production queries after re-clustering on(date, hour, source, id)from partition-by(date, hour)). OPTIMIZE engineering disclosure: planning phase 12h → 23m on 10 PB tables (31×); execution phase 5× faster on Medium DBSQL clusters — the engineering work behind Liquid Clustering's PB-scale viability. Forward-looking: co-clustered joins (Private Preview, "~51% faster, 87% less shuffle data" on Liquid-to-Liquid joins) + in-place Liquid Conversion (Private Preview SQL surface disclosed). Existing pages extended: systems/liquid-clustering gains comprehensive Eight myths debunked + Production case studies at PB scale + OPTIMIZE engineering improvements + Forward: co-clustered joins sections; systems/delta-lake gains a twelfth face (transaction-log-based-pruning) as the load-bearing architectural fact; systems/databricks-predictive-optimization gains the PB-scale role disclosure (Arctic Wolf attribution); automatic-table-optimization gains the OPTIMIZE engineering improvements; concepts/write-amplification gains Z-Order's periodic-rewrite cost geometry as a wiki-canonical instance at the lakehouse table-layout altitude. Caveats: consultative-listicle framing for a Tier-3 source; the 75% over-partitioning rate is unscoped (corpus / methodology not disclosed); benchmark numbers (22% / 35%) cite "a real-world data warehousing benchmark" without naming the benchmark; Arctic Wolf attribution between Liquid Clustering / Predictive Optimization / UC managed tables / Delta is not separately quantified; co-clustered joins and in-place Liquid Conversion are Private Preview without GA timelines; Z-Order critique is asymmetric (post elides cases where Z-Order remains workable); no QPS / concurrency / lock-contention numbers for the named customers. Frontmatteringested:flippedfalse → true. -
2026-05-29 — sources/2026-05-29-databricks-enabling-evolutionary-database-development-database-branching-with-lakebase (Tier-3 Databricks Blog — Part 1 of a three-part Evolutionary Database Development series; passes scope despite narrative- driven shape because Tier-3 source names a real production substrate, frames Practice #4 as the constraint that the substrate lifts, names a public open-source IDE extension, and discloses the four-validation CI flow). Enabling Evolutionary Database Development: database branching with Lakebase. First wiki canonicalisation of evolutionary-database-design as a discipline — Martin Fowler's 2003 essay → Pramod Sadalage's 2006 Refactoring Databases with 70+ named refactorings → Humble & Farley's 2010 Continuous Delivery Chapter 12 ("Managing Data") → 2026 Lakebase substrate change. Load-bearing methodology argument:
Practice #4 — "everybody gets their own database instance"
has stayed aspirational for twenty years because per-developer
production-shaped databases cost time, money, and DBA cycles;
the
compensating layer that emerged (mock objects, in-memory DB
substitutes like H2/SQLite, shared staging environments, DBA
ticket queues) "became foundational methodology by default,
not by design." Verbatim canonicalising claim: "In 2026,
copy-on-write database branching arrives in Databricks
Lakebase. A one-second, zero-storage-at-creation branch of a
terabyte-scale production database is now an O(1) operation.
The constraint that kept Practice #4 aspirational has lifted."
Three load-bearing properties of the developer-DB instance
canonicalised: fast (created when needed) + realistic
(same Postgres engine, same governance, production-shaped
data) + isolated (experiments don't interrupt anyone) —
each historical compensating-layer alternative violates at
least one. Re-uses Fowler's 2003 protagonist Jen + the
Split Column refactoring
as the worked example — "Same Jen. Same refactoring. What
changed is the capability." CI flow canonicalised verbatim:
"CI does what Jen just did, but for the team: it creates its
own temporary Lakebase branch, applies the migration, runs the
application test suite, runs database tests against the
migrated schema, validates the migration itself (applies
cleanly, idempotent, reversible), and posts a schema-diff
comment on the PR showing exactly which database objects
changed." Four-validation bundle (applies cleanly + idempotent
+ reversible + application tests) is the substrate change that
absorbs the breakage-class question previously held by the DBA.
DBA reframe canonicalised as the role-evolution payoff:
"the DBA can review on their schedule, not Jen's... improve
the solution around data integrity, indexing strategy, future
extensibility or long-term maintainability, not on the
protective gatekeeping that used to take all their time."
Migration tools cited as platform-agnostic: Flyway,
Liquibase, Alembic, Knex, Prisma — substrate change is
orthogonal to the tool ecosystem. New named system:
systems/lakebase-scm-extension — public open-source
VS Code / Cursor IDE extension at
github.com/databricks-solutions/lakebase-scm-extension
that synchronises a developer's git branch with a matching
Lakebase database branch and surfaces the Branch Diff
Summary view. Forward references: Part 2 (Jen's New
Playbook — copy-on-write internals + methodology
optimisations); Part 3 (Jen's Team at Scale — 50-developer
governance + agent-in-the-loop + DBA re-deployment); Companion:
Plugin Walkthrough (Lakebase SCM Extension end-to-end);
Lakebase App Dev Kit for agents with companion ebook. New
wiki entities: 4 concepts (concepts/evolutionary-database-design,
practice-4-everybody-gets-their-own-database-instance,
database-development-compensating-layer,
dba-as-design-collaborator); 3 patterns
(patterns/per-developer-database-branch-paired-with-code-branch,
patterns/database-branch-per-test-over-mocking,
migration-script-travels-with-application-code);
1 system (systems/lakebase-scm-extension). Existing
pages extended: systems/lakebase gains the
evolutionary-database-development capability slice as the
newest section under "Capabilities surfaced so far";
concepts/database-branching gains a fifth canonical
use-case axis (after PlanetScale schema-change, LangGuard
governance-policy-testing, Stripe Projects agent-operation,
Backstage state-heavy-application IDE/CI/QA) — methodology-
substrate fit; concepts/copy-on-write-storage-fork gains a
fifth canonical instance with the methodology arc;
patterns/database-branch-per-test-over-mocking gains a
second canonical instance pairing the mock-replacement axis to
the Practice #4 substrate-shift framing. Sibling-cluster
cross-refs: positioned as methodology-arc complement to
the
Backstage POC (substrate disclosure) + the
Backstage Part 2 (governance composition) — same Lakebase
branching primitive, three different framings (technical
POC; governance composition; methodology arc). Sibling to
LangGuard (governance-policy-testing axis) and
Stripe Projects (agent-operation axis) — the methodology arc
unifies all three under "what becomes operational when Practice
#4 is finally affordable." Sibling to
PlanetScale's branch-based schema-change workflow as the
earlier production-realisation of the same constraint-lift
(PlanetScale 2021 vs Databricks 2026); PlanetScale puts the
branching at the deploy-time layer with a deploy-request +
queue + traffic-aware-throttler, the Lakebase flow puts it at
the IDE / CI / per-PR layer with same-PR migration scripts.
Caveats: narrative-driven shape; no new architectural
disclosure beyond prior Lakebase posts (the methodology arc is
the contribution); no SCM Extension internals; no multi-
developer scaling disclosure (deferred to Part 3); no agent-
substrate detail (deferred to App Dev Kit); no quantitative
cost / billing detail beyond "zero storage at creation"; tool
list (Flyway/Liquibase/Alembic/Knex/Prisma) is platform-
agnostic by design but means engineers don't get tool-selection
guidance. Frontmatter ingested: flipped false → true.
wiki/index.md auto-regenerated by build watcher (Distilled
+1; Concepts +4; Systems +1; Patterns +3; Companies 40
unchanged; Latest additions list will prepend the new source).
- 2026-05-29 — sources/2026-05-29-databricks-databricks-at-sigmod-2026
(Tier-3 Databricks Blog conference-announcement post — passes
scope despite announcement-shape because each named architectural
claim is real and consequential, with arXiv-level paper backing).
Databricks at SIGMOD 2026. Short corporate-blog post that
nevertheless discloses the first publicly named architecture of
Databricks' incremental-view-maintenance
engine — Enzyme — which powers the
materialized-view track of Spark
Declarative Pipelines (SDP). Two papers announced: SIGMOD
2026 honorable-mention "Enzyme: Incremental View Maintenance
for Data Engineering"
(arXiv:2603.27775; presented
by Ritwik Yadav) and VLDB 2026 "A Decade of Apache Spark
Structured Streaming: How We Evolved the Architecture To Meet
Real-world Needs". First wiki disclosure of the SDP two-track
model: "There are two ways to write incremental programs in
Spark Declarative Pipelines (SDP), and customers can mix-and-match
these within a pipeline." (a) Enzyme for
declarative
@dp.materialized_view; (b) Structured Streaming for explicit stateful operators / watermarks / custom aggregations. MV-as-ETL thesis: "Materialized views (MVs) are popular for query acceleration… When creating SDP, we decided to go beyond query acceleration and apply materialized views to the extract-transform-load (ETL) use cases. Our key observation is that if MVs can be efficiently and incrementally maintained, it will significantly simplify ETL workloads which otherwise require writing complex custom code." Enzyme's four novel claims over prior industrial IVM: (1) full MV-grammar coverage — joins -
window functions + aggregations + combinations, all incrementally maintained in one engine; (2) non-deterministic functions —
current_date()and AI functions handled correctly under incremental maintenance, where most prior industrial IVM rejects them or recomputes in full; (3) multi-language MVs — Python + SQL, with change detection on Python MV definitions as a named open problem solved; (4) cost-model-driven incrementalisation — runtime choice between partition-level vs row-level updates, selective intermediate-result caching, cost model fed by plan information + prior execution statistics. Performance disclosure: relative speedup chart vs "another competing industry solution (name anonymized to CV-IVM due to licensing restrictions)"; absolute numbers / workload axes / ablations deferred to the paper. Bangalore named as "a large Databricks R&D hub" and SIGMOD 2026 host city. New systems (1): systems/enzyme-ivm (also disambiguates against the unrelated Airbnb React testing utility of the same name). New concepts (4): concepts/incremental-view-maintenance (parent technique, full page); concepts/materialized-view (foundational concept, full page); non-deterministic-mv-maintenance (the hardest IVM challenge, full page); multi-language-materialized-view (Python + SQL MVs, full page). New patterns (1): cost-model-driven-incrementalization-strategy (sibling to Spark AQE's plan-rewriting shape, applied to IVM rather than query execution). Extended pages (3): SDP gains a two-track-architecture section + new "Seen in" entry (now the third source, the canonical architecture-level reference); Structured Streaming gains VLDB 2026 forward-reference + new "Seen in" entry (first wiki source naming Structured Streaming's academic publication); Apache Spark gains an academic-publication face noting the two-paper SIGMOD/VLDB pair as the first wiki disclosure that modern Databricks Spark contributions are published as named academic artefacts at top systems venues. Sibling-cluster cross-refs: complements prior SDP-naming sources (multimodal, AutoCDC) by revealing the engine inside the decorator — those sources used@dp.materialized_viewas a black box; this source names the IVM engine and its capabilities. Sibling to Octopus Energy MHHS on the trust-the-optimiser axis: AQE for query execution there, Enzyme cost-model for IVM strategy here. Caveats: announcement-shape, not architecture deep-dive; no absolute numbers (the only chart is relative speedup vs anonymised CV-IVM); mechanism for non-deterministic-function correctness not disclosed; Python-MV change-detection technique not disclosed; cost-model feature set not enumerated; relationship to Catalyst not described; workload-class boundaries (≥3-table joins, recursive CTEs, foreign vs managed Iceberg) not characterised; CV-IVM identity withheld under licensing. Frontmatter: rawingested: false → true. Return contract:ingested: wiki/sources/2026-05-29-databricks-databricks-at-sigmod-2026.md. -
2026-05-28 — sources/2026-05-28-databricks-advancing-apache-iceberg-on-databricks-iceberg-v3-ga-open-sharing-and-unified-governance (Tier-3 Databricks Blog feature-roundup post — passes scope on shape-canonicalisation grounds despite the marketing-roundup framing because each named primitive is a real architectural element with consequences for the OTF ecosystem). Advancing Apache Iceberg on Databricks: Iceberg v3 GA, Open Sharing, and Unified Governance. Coordinated set of Iceberg capability releases that reposition Unity Catalog as a fully Iceberg- native catalog. Format-level GA: Iceberg v3 reaches GA on Databricks across managed / foreign / UniForm-enabled tables — three primitives ship together: deletion vectors (file-level row- delete representation; merge-on-read applied to deletes), row tracking (stable per-row identity supporting more efficient incremental processing), VARIANT type (standard representation for semi-structured data). All three already existed on the Delta side; v3 brings parity, with cross-format compatibility called out verbatim. Catalog-side primitives: Managed Iceberg (GA) — UC creates / reads / writes / governs Iceberg tables directly with Predictive Optimization + Liquid Clustering applying. Foreign Iceberg (GA) + Credential Vending for Foreign Iceberg (GA) — UC governs Iceberg tables managed in eight named external catalogs (AWS Glue, Snowflake Horizon, Hive Metastore, Apache Polaris, Salesforce Data Cloud, Google Cloud Lakehouse, Palantir, Workday) while leaving data in place; mints short-lived scoped credentials. External Sharing to Iceberg clients (GA) — Delta Sharing now emits Iceberg REST endpoints; recipients on Snowflake / Trino / Flink / Spark consume shared data via Iceberg-compatible clients without ingestion or copies. Plus External Sharing of Foreign Iceberg tables (Public Preview). Cross-engine ABAC (Beta) — UC ABAC policies evaluate during server-side scan planning via the Iceberg REST Catalog Scan Planning API (Iceberg 1.11). The catalog returns a filtered scan plan; the engine reads only authorised data. Compatible engines: Spark / DuckDB / any engine implementing the Iceberg-1.11 scan-planning client. Wiki's first canonical instance of scan-planning-as-policy-enforcement-point — extends UC ABAC beyond the Databricks-compute boundary. Iceberg-compatible materialized views (Gated Public Preview) — managed MVs exposed downstream as native Iceberg tables; syntax
CREATE MATERIALIZED VIEW my_mv USING ICEBERG. Forward-looking: Iceberg v4 + Delta 5.0 alignment on a shared "adaptive metadata tree" metadata structure (format-co-evolution-iceberg-delta) — wiki's first canonical disclosure of explicit OTF-format-convergence direction. New systems (2): systems/iceberg-v3 (canonical v3 milestone page); systems/iceberg-rest-catalog-scan-api (the Iceberg-1.11 scan-planning client surface). New concepts (5): concepts/tombstone, row-tracking, variant-type, concepts/attribute-based-access-control, foreign-iceberg-table, format-co-evolution-iceberg-delta. New patterns (1): scan-planning-as-policy-enforcement-point. Extended (5 existing pages): systems/apache-iceberg (gains Iceberg v3 GA section + catalog-side-primitives section + Seen-in entry — 15th Iceberg face); systems/unity-catalog (gains 6th face — fully-Iceberg-native catalog with five concurrent surface-area expansions); systems/delta-sharing (gains bi-format recipient surface section + Seen-in entry — recipients on Iceberg clients now consume the protocol); systems/delta-lake (gains 11th face — format-co-evolution- with-Iceberg, with verbatim Delta 5.0 disclosure and cross- format-compat for v3 features); systems/unity-catalog-abac (gains cross-engine-ABAC Beta Seen-in entry — first wiki disclosure of UC ABAC moving beyond Databricks compute); concepts/attribute-based-access-control (third canonical instance — cross-engine axis); concepts/short-lived-credential-auth (gains foreign-Iceberg face — cross-catalog credential vending). Tier-3 marketing-roundup framing acknowledged throughout; no quantitative numbers anywhere in the source; mechanism depth on any single primitive deferred to spec / docs. Architectural significance is consolidation of named entities under one catalog-vendor surface area, not deep mechanism disclosure. -
2026-05-27 — sources/2026-05-27-databricks-reliable-llm-inference-at-scale (Tier-3 Databricks inference platform team — Marius Seritan, Cyrielle Simeone, Andy Zhang, Yu Zhang, Nick Lanham — passes scope on distributed-systems-internals / production-LLM-serving- architecture / multi-tenant-capacity-management grounds). Reliable LLM Inference at Scale. Production architecture behind Databricks' multi-tenant LLM-serving platform processing 125T+ tokens/month of frontier-model traffic. Six structural contributions: (1) Axon named publicly for the first time — the LLM data-plane router built on Dicer — sitting between rate-limiting and the inference runtime. (2) Model units as the LLM-request-cost abstraction: `cost ≈ α·input + β·output
-
γ·...` with β > α (decode > prefill), coefficients per (model, hardware) via auto-benchmarking, prefix-caching + multi-modality modifiers. The unit-of-account that lets Databricks offer VM-equivalent capacity guarantees instead of best-effort capacity. (3) Cost-based load balancing via patterns/ai-gateway-provider-abstraction — Dicer-keyed-on-MUs replaces P2C-with-active-requests at LLM scale (verbatim retire-P2C datum: "LLM latencies tend to be high, server counts are lower than scaled out CPU systems, and the cost of misrouting is severe"). (4) Cost-based autoscaling via model-units-utilization-autoscaling — MU utilisation ratio averaged across pods drives a model-agnostic scaling infrastructure. >80% GPU savings vs static-peak provisioning on bursty workloads. (5) Stateful sessions via stateful-llm-session-routing — workload requests pin to a Dicer-assigned pod subset for KV prefix-cache locality + bounded blast radius (two purposes, one primitive). (6) Runtime reliability via two named mechanisms: silent-hang detection through prioritised black-box health checks (highest scheduling priority, <5-min detect→kill→recover cycle, false liveness-probe failures several/week → zero); and the multimodal CPU bottleneck fixes — Torchvision over PIL (10× preprocessing speedup) + OMP_NUM_THREADS fix (avoid container CPU throttling from thread oversubscription, e.g. 192 host vCPUs → 12 container vCPUs). Combined: >3× RPS jump on same hardware. Customers named: Superhuman, YipitData, Fox Sports. Hosted models: frontier OS (Kimi, Qwen) + proprietary (OpenAI, Gemini, Claude). New systems (1): systems/databricks-axon. New concepts (6): model-units, concepts/model-flops-utilization, silent-hang-llm-server, multimodal-cpu-bottleneck, omp-num-threads-container-misconfiguration, multi-tenant-llm-capacity-allocation, non-uniform-llm-request-cost. New patterns (5): patterns/ai-gateway-provider-abstraction, model-units-utilization-autoscaling, prioritized-black-box-health-check, torchvision-over-pil-image-processing, stateful-llm-session-routing. Extended: systems/dicer gains a third canonical Databricks face (LLM-router-substrate); systems/databricks-model-serving gains an LLM-specific architecture section structurally distinct from the EDS+P2C+request-concurrency Superhuman face; concepts/power-of-two-choices gains its first canonical retire-for-LLM-serving datum; request-concurrency-as-autoscaling-signal gains its next-evolution-to-MUs cross-reference. The platform's reliability story is now documented at platform-internals depth across two layers: the platform/runtime layer from the Superhuman post (2026-05-08) and the LLM-router/multi-tenant-capacity layer from this post.
-
2026-05-27 — sources/2026-05-27-databricks-how-the-lakebase-architecture-stays-resilient-to-cloud-failures (Tier-3 Databricks Engineering reliability roadmap on Lakebase / Neon — passes scope on distributed-systems-internals / production-AZ-outage / serverless-Postgres-SLO grounds). Authors: Jasraj Dange, Hans Norheim, Stas Kelvich, John Spray. Reframes serverless-Postgres reliability for the agentic-workload era — "agents create 4× as many databases as humans do"; "starting tens of millions of databases every day". Six pillars: (1) stateless Postgres compute + zone-redundant storage default-on for all tiers — eliminates the hot-standby tax + 10s- of-minutes WAL-replay crash recovery (verbatim); (2) "control plane is the new data plane" — split a hot-path data-plane controller off the management plane; empirical signal: 90% of compute sessions for auto-suspending databases in Neon are <10 min; (3) Critical- path dependency minimisation via bare- metal pool + own vertical-autoscaling virtualisation layer + own zone-resilient storage — collapses the 5-link cloud- provider control-plane chain to one already-completed dependency; (4) Cell-based architecture with regional cell composition — canonical production-AZ- outage instance on 2026-05-08 us-east-1 thermal-event: "the cell-based architecture reduced the impact by roughly an order of magnitude" — ~13% of databases in the region affected (1/8 = 8 cells, 7 failed-over correctly + 1 imperfectly); (5) Failure simulation + injection with failpoints + SQLancer / SQLsmith correctness validators, escalating to whole-AZ network-partition drills with 30-second-or-better per-database outage target; (6) Per-database availability attainment as the SLO measurement substrate (vs fleet-aggregate); two-bar reporting (99.95% / 99.99%); disclosed 2026 H1 attainment table; five SLI menu including the serverless-specific database startup time. New systems (2): systems/neon (first dedicated page), systems/sqlsmith (first wiki disclosure). New concepts (6): concepts/control-plane-data-plane-separation (the central reframe), zone-redundant-storage (the storage-tier property), critical-path-dependency-minimization (the cloud-provider-control-plane-bypass discipline), database-availability-attainment (per-database SLO shape), database-startup-time-sli (the serverless-specific SLI), concepts/chaos-engineering (the next-level drill), concepts/chaos-engineering (the in-code injection primitive). New patterns (4): preallocated-bare-metal-pool-with-virtualization (the cloud-provider-control-plane-bypass primitive), separate-data-plane-controller-for-hot-path (the control-plane decomposition pattern), per-database-availability-attainment (the SLO measurement pattern), whole-az-network-partition-drill (the chaos drill). Extended (10+ existing pages): systems/lakebase (new reliability-roadmap section + frontmatter), systems/pageserver-safekeeper (new zone-redundant-storage disclosure), systems/sqlancer (first canonical explicitly-adopted instance), concepts/blast-radius (~13% / ~order-of-magnitude quantified production-AZ-outage instance), concepts/cell-based-architecture (canonical serverless-database regional-composition instance), concepts/control-plane-data-plane-separation (inversion- corner-case under agentic workloads), concepts/static-stability (fourth canonical instance at cloud-provider-control-plane-bypass altitude), concepts/stateless-compute (Postgres stateless-compute reliability framing), concepts/chaos-engineering (Lakebase three-altitude regime), concepts/chaos-engineering (sibling whole-AZ-partition shape), patterns/cell-based-architecture-for-blast-radius-reduction, continuous-fault-injection-in-production. Caveats: data-plane controller and whole-AZ partition drill are "in flight" not landed; cell-count for us-east-1 implicit not stated; vertically-autoscaling virtualisation layer linked but not detailed; April-2026 attainment dip (99.96 → 99.93; 99.81 → 99.75) unexplained.
-
2026-05-27 — sources/2026-05-27-databricks-bi-serving-pointers-maximizing-for-performance-and-tco (Tier-3 Databricks Engineering BI-serving walkthrough — passes scope on the substantive architectural-framings ground despite the vendor-walkthrough format). Frames the four-layer BI serving stack on the Databricks Lakehouse (physical storage → semantic layer → automatic materialization → DBSQL warehouse / caching tier) with the thesis that "each layer compounds the performance gains of the layer below it". Six new wiki pages:
- systems/databricks-metric-views (first wiki disclosure
as a distinct named primitive) — UC-resident
headless-BI semantic layer; define metrics ONCE, every consumer
(AI/BI Dashboards / Genie / SQL notebooks / third-party BI
tools) resolves the same
MEASURE()definition; semantic metadata (display_name/comment/synonyms) is the AI-grounding contract for Genie. - systems/databricks-predictive-optimization (first
wiki disclosure as a dedicated system page) — auto-
OPTIMIZE/VACUUM/ stats collection inline-during-Photon-writes plus existing-table back-fill; 22% average performance improvement in observed workloads; newCLUSTER BY AUTOextension to Liquid Clustering. - systems/databricks-sql-warehouses (first wiki
disclosure as a distinct system) — serverless auto-scaling;
two-tier cache hierarchy (disk cache + Query Result
Cache) for repetitive BI workloads; reflexive
system.billing.usage/system.query.historyobservability surface. - systems/databricks-ai-bi-dashboards (first wiki disclosure) — first-party Databricks dashboard surface; one of four named consumers of Metric Views.
- headless-bi-semantic-layer (first wiki canonicalisation) — "define metric ONCE, every consumer resolves the same"; AI-grounding via semantic metadata sibling to layered grounded context from Cloudflare Skipper.
- metric-view-materialization (first wiki
canonicalisation) — automatic pre-aggregation + incremental
refresh + intelligent query rewriting + transparent
routing; collapses three coupled artifacts (aggregate tables
- refresh pipelines + BI-tool query updates) into one governed primitive.
Two new concepts: automatic-table-optimization (the substrate-owned-OPTIMIZE / VACUUM / stats shape, beyond Databricks-specific framing); dbsql-caching-tiers (the two-tier disk-cache + QRC hierarchy); optimizer-statistics-as-skipping-substrate (the reframing that stats are not just plan-quality input but the substrate that makes data skipping possible).
Four new patterns: governed-metric-as-headless-bi-substrate (define metrics in the catalog, every consumer resolves the same); auto-materialized-aggregation-via-semantic-layer (enable materialization on the metric, no separate aggregate tables / refresh pipelines / BI-tool queries to maintain); managed-table-as-default-storage-layer (use UC managed tables across all medallion layers, not just Gold); query-rewrite-to-pre-aggregated-materialization (the optimiser-side counterpart pattern that makes transparent routing work).
Existing pages extended:
- systems/liquid-clustering gains the CLUSTER BY AUTO
Predictive-Optimization-driven automatic key-selection
disclosure + the "replaces static partitioning AND manual
Z-ORDER, redefinable without rewriting existing data"
framing.
- systems/uc-managed-tables gains the "foundation of
everything in the BI serving stack" framing + the
cross-medallion-layer recommendation.
- systems/unity-catalog gains a twelfth canonical face
— BI-serving / semantic-layer substrate hosting Metric
Views.
- systems/databricks-genie gains a Metric-Views consumer
role with semantic-metadata-as-prompt-engineering grounding.
- concepts/oltp-vs-olap gains a Gold-tier BI-serving
canonicalisation alongside its prior LLM-pipeline-state
canonicalisation; names the platform-resident dimensional-
modelling primitives (PK / FK with RELY hint, identity
columns, CHECK / NOT NULL) and the Silver-Data-Vault →
Gold-star-schema layering recommendation.
Operational disclosures: the only quantitative figure is the 22% average performance improvement from Predictive Optimization in observed workloads (cited via a separate Databricks post). The article is otherwise architectural — no concurrency / scaling envelope, no per-tenant numbers, no multi-tenant isolation depth, no materialization freshness contract under high ingest, no query-rewriter coverage envelope. Open standard provenance: SPARK-54119 (Apache Spark Metric Views OSS implementation) + UC OSS support coming. Closing thesis: "Each layer compounds the performance gains of the layer below it." The compounding shape — physical-layer optimisation → fewer rows scanned at the materialization layer → fewer rows aggregated at the semantic layer → faster consumer queries — is the architectural payoff over per-tool BI semantic layers + hand- built aggregate tables + refresh pipelines.
Sibling to Cloudflare Town Lake / Skipper (both arrive at the same architectural thesis from opposite directions — Cloudflare's R2 Data Catalog composes managed-Iceberg-on-R2 + DataHub + Trino + Skipper at the data-platform layer; Databricks composes UC managed tables + Metric Views + DBSQL warehouses + Genie at the BI-serving layer; both expose semantic metadata as the AI-grounding substrate). Sibling to sources/2026-05-14-databricks-expanded-interoperability-with-unity-catalog-open-apis (both name UC managed tables as the foundation; the May-14 post is open-engine-access-focused, the May-27 post is BI-serving-stack-focused). - 2026-05-23 — sources/2026-05-23-databricks-scaling-for-mhhs-octopus-energy-50x-cost-reduction (Tier-3 Databricks customer-success post on the Octopus Energy margin data pipeline rebuild for the UK regulatory transition to Market-wide Half-Hourly Settlement — passes scope on production-architecture- internals + scaling-trade-offs grounds despite the customer-success framing. Volume-driver: 2 meter reads/customer/month → 48 reads/customer/day = 48× data-point increase for 8M+ customers, projecting +$1M/yr in unsustainable compute under the legacy monolithic-monthly-grain pipeline. The architectural diagnosis is the article's central concept — grain misalignment: "The legacy pipeline had been built around a single grain: monthly. […] Running all three through a single monolithic pipeline meant processing the entire dataset on every run, regardless of what had actually changed." The rebuild splits the monolith into three grain-aligned streams (Settlement HH for industry settlement / Half-Hourly for smart-tariff revenue (EVs, heat pumps, ToU) / Monthly for standard tariff) on a unified multi-terabyte multi-grain source-of-truth layer — orchestrated by a "Job of Jobs" parent-child pattern preserving each stream's independent tuning profile ("what works as a Spark optimisation for Settlement is not necessarily right for NHH"). The single highest-leverage optimisation: Delta CDF-based incremental processing of the upstream consumption layer — rows/run dropped 25 B → 300 M (98.8% reduction), freshness weekly → daily. Layered Spark/Delta optimisations: broadcast joins for reference tables under 500 MB (eliminating shuffle on multi-key joins with date ranges), liquid clustering on filter/join columns ("avoids the small-file problem, higher memory consumption, and I/O overhead that come from over-partitioning"), lineage simplification + early column / row pruning, and the counter-instinct fourth lever — removing custom optimisation code in favour of Spark AQE ("In several cases, Spark's Adaptive Query Execution (AQE) outperformed hand-tuned logic. The team removed custom optimisation code and let AQE do its job"). Databricks Serverless named as the development-velocity enabler — "The testing and development process could not have been done without serverless. Using the serverless UI helped us to identify bottlenecks and make easy comparisons between different runs" — zero cluster startup + side-by-side run comparison made the three-month delivery window viable. Final geometry: $0.48 per settlement date — 50× below the projected MHHS cost ($23.63), 2× below the legacy ($0.71) despite processing 48× more data points. ~$1M annualised cost avoidance excludes the upstream-incremental savings, which are additional. Three engineers, three months. Generalisation made explicit: "Any time a system moves from monthly to daily, daily to real-time, or aggregate to transactional, the same dynamics apply" — four transferable takeaways: grain misalignment is the hidden cost driver, incremental processing transforms pipeline economics, remove before you add, trust the optimiser. New systems (3): Octopus Margin Data Pipeline, Spark AQE (promoted from recurring-tag mention to dedicated page), Liquid Clustering (promoted from recurring-tag mention to dedicated page). New concepts (3): grain-misalignment, data-pipeline-grain, remove-before-add-optimization. New patterns (4): grain-aligned-stream-split, cdf-incremental-replacing-full-rescan, broadcast-join-for-small-reference-tables, job-of-jobs-orchestration. New company: Octopus Energy — UK retail energy supplier, 8M+ customers; first wiki disclosure. Extended: concepts/change-data-capture gains its multi-terabyte upstream-substrate face (distinct from the Bronze→Silver promotion face from the Claroty source); systems/delta-lake, systems/apache-spark, systems/databricks-serverless-compute each gain a new "seen in" face. Saad Ali, Lead of the Margin Data Team at Octopus Energy: "You can't just throw more compute at a problem like this. You have to rebuild and rethink your logic from the ground up.").
- 2026-05-22 — sources/2026-05-22-databricks-observability-any-agent-anywhere-otel-unity-catalog
(Tier-3 Databricks blog GA announcement of OTel-format trace
ingestion direct to Unity Catalog Delta
tables — the agent-side OTel companion to the 2026-05-20
full-payload Inference Tables
substrate. Three new systems canonicalised:
Zerobus Ingest (managed serverless
OTLP/gRPC + REST receiver — "With a 'single-sink' architecture,
Zerobus Ingest simplifies observability by streaming data
directly to the lakehouse. Existing OLTP-compatible collectors
can point directly to this endpoint via gRPC, entirely bypassing
intermediate message buses like Kafka"),
UC OTel Trace Tables (six MLflow-
derived Delta views:
<prefix>_otel_spans,_otel_logs,_otel_metrics,_otel_annotations,_trace_unified,_trace_metadata— auto-liquid-clustered, MLflow per-experiment trace cap removed, unbounded storage), and MLflow OTel Tracing (the framework-side instrumentation surface —mlflow.<lib>.autolog() -
@MLflow.tracedecorator combined; provisions the UC tables from Python; agent-runs-anywhere portability "In fact the support assistant agent example that was used for this blog is deployed locally"). Three new concepts: single-sink-telemetry-architecture (the no-broker shape — managed receiver direct to Delta), concepts/observability (OTel as protocol-portable boundary — "using the OTel standard to separate instrumentation from storage"), and production-traces-as-evaluation-substrate (durable prod traces become MLflow eval-dataset bootstrap — "these prompts originate from actual user interactions, they better represent the scenarios your agent must handle compared to purely synthetic test cases" — same judges run continuously on live traces). Three new patterns: patterns/telemetry-to-lakehouse (the Zerobus shape canonicalised), bootstrap-eval-dataset-from-production-traces (SQL-warehouse-driven dataset materialisation), and component-level-latency-from-otel-spans (per-tool P50/P99 dashboards over_otel_spans, finer than trace-level latency that native dashboards default to — "That tells us whether the LLM, a Genie tool call, or another step is the bottleneck"). Operational disclosures: 200 QPS starting ingest throughput (account-team escalation for higher), storage limit none, MLflow per-experiment trace cap removed, auto liquid-clustering post latest product update. Three SaaS-vs-lakehouse asymmetries argued verbatim: "retention economics" (object storage cheaper than SaaS), "the PII deadlock" (no third-party data egress — UC column masking + row filtering apply automatically because traces are governed Delta tables), and "analytics, not just telemetry" (joinable with business data). Customer scale points named: Experian (Eva + Latte agents — "hundreds of thousands of traces"), Superhuman/Grammarly ("hundreds of thousands of traces per day" — explicitly replacing a custom point solution: "that maintenance burden was a real pain point for our teams"), SmartSheet (two production agents in "three-day co-build", "tens of thousands of evaluations"), The Standard (insurance underwriting + claims agents). Reference agent: LangGraph + Databricks-hosted Claude Sonnet 4.6 + Genie tool over MCP, deployed locally. Cross-links extended: patterns/telemetry-to-lakehouse gains its third citation — agent-trace specialisation distinct from the 2026-04-17 metrics+traces face and the 2026-05-20 full-payload face; systems/mlflow gains a sixth+ face (OTel-tracing-direct-to- UC, with the prod-traces-as-eval-substrate flow); systems/opentelemetry gains the "protocol-portable boundary between agent instrumentation and lakehouse storage" face; concepts/observability extends from the 2026-05-05 metric-time-series face to span/log/metric agent- trace face; systems/inference-tables gains a sibling- substrate cross-reference distinguishing model-call-payload granularity (Inference Tables) from agent-execution-span granularity (UC OTel Trace Tables) — both UC-resident, both governed under one catalog.) -
2026-05-22 — sources/2026-05-22-databricks-how-world-bank-group-uses-databricks-to-eradicate-poverty-through-shared-knowledge (Tier-3 Databricks customer-success post on World Bank Group's Knowledge 360 + Data 360 unified-platform build. Architectural primitives: per-domain Genie instances each pinned to a metrics layer + RAG agent over UC Volumes + Vector Search + agentic router with three named classifiers (intent / domain / decomposer)
-
decoupled visualisation agent + AI Gateway as control plane. Two new architectural disclosures: (1) Genie's default LLM-only output is nondeterministic enough to be unfit for financial / operational reporting (Suresh Kaudi: "In the structured content, you need an answer. What is my bank balance? I don't want to see a different number every time") — drives the metrics-layer-for-deterministic-Genie-answers retrofit. (2) Each Genie is per-metrics-layer / per-domain, so cross- domain queries break single-Genie — drives the intent-domain- decomposer agentic-router fan-out shape. Operational scale: 3M document downloads / month through the AI-powered layer, half from low- and middle-income countries; external-feedback prototype built and deployed in ~2.5 days. Borderline-include Tier-3 ingest — mechanism-light throughout (classifier model choice, decomposition strategy, metrics-layer implementation substrate not disclosed) but the architectural shape and failure modes are specific enough to canonicalise. New patterns (2): intent-domain-decomposer-agentic-router (the three-classifier router with set-output-and-fan-out generalisation of multi-agent supervisor routing); metrics-layer-for-deterministic-genie-answers (the pin-Genie-to-metrics-layer retrofit for same-question-same-answer contracts). New company: companies/world-bank-group (canonical wiki home for the WBG knowledge platform; complements the existing Tier-3 Databricks-customer corpus). Cross-links added: sixth canonical face of Genie on the wiki; first Genie deployment fronted by an intent-domain-decomposer router. Composes with patterns/specialized-agent-decomposition (Virtue Foundation's alternative-selection sibling pattern) and patterns/upstream-the-fix (Trinity Industries' upstream measure-consolidation complement). Confirms the 2026-05-20 Governing AI agents at scale scope-generalisation thesis (AI Gateway gates non-coding-agent populations too).)
-
2026-05-22 — sources/2026-05-22-databricks-accelerating-llm-inference-with-prompt-caching-for-open-source-models (Tier-3 Databricks blog post; short GA announcement of implicit prompt caching for open-weights models served on the Foundation Model APIs). Architectural news in three load-bearing design choices: (1) implicit — "customers do not need to configure anything, our system has built to automatically run the prompt caching and reuse to improve throughput"; (2) volatile-only, tenant-isolated, never persisted — the safety envelope that lets default-on caching ship on multi-tenant infrastructure without an encryption-at-rest threat model; (3) inherited platform-wide — Agent Bricks, Genie, AI Functions all inherit caching at no integration cost. Disclosed numbers from the GPT-OSS production batch-inference rollout: +2.5× per- replica input-token throughput, 3× P50 latency reduction at a 30% cache hit ratio. Catalog covered: GPT-OSS 20B + 120B, Gemma 3 12B, Llama 3.1 8B (including PEFT-served fine-tuned variants), Llama 3.3 70B. New system (1): Databricks FMAPI Prompt Caching — the named platform feature. New concepts (2): implicit-prompt-caching (the design choice that caching is platform-decided, zero-configuration, and contrasts with explicit caching APIs from Anthropic / OpenAI / Google); volatile-only-prompt-cache-isolation (the multi-tenant security shape composing tenant isolation + RAM-only residency + no persistence). New pattern (1): implicit-prompt-cache-as-platform-default (the default-on substrate-layer rollout pattern, contrasted with explicit-API integration). Cross-link added: companion to concepts/context-engineering (Cloudflare's contrasting client-signal-driven design) and concepts/kv-cache (the underlying primitive being reused).
-
2026-05-20 — sources/2026-05-20-databricks-marketing-campaigns-with-lakebase (Tier-3 Databricks blog post; integration tutorial for SAP Engagement Cloud at Deichmann + architecture pitch for Lakebase as the canonical bursty-OLTP backend for marketing campaigns). Borderline-include: ~70% click-through tutorial, ~30% architecture, but the architecture content surfaces three genuinely-new wiki canonicalisations. New systems (2): Lakebase Synced Tables (managed Delta → Postgres materialisation with three sync modes — snapshot / triggered / continuous — and the load-bearing operational rule "when more than 10% of the data is updated, we recommend snapshot mode, which delivers 10x better performance than triggered mode"); Lakehouse Sync (the Postgres → Delta direction — "a native, continuous CDC-based pipeline from Lakebase Postgres to Unity Catalog Delta tables" — closing the bidirectional governed-data path between operational and analytical tiers without hand-maintained CDC pipelines). New concept (1): Lakebase Local File Cache (LFC) — first wiki disclosure of LFC as Lakebase's compute-VM-local cache of Pageserver pages, alongside the two Lakebase-specific Postgres query statistics
PREFETCH(prefetch requests issued/hit/wasted) andFILECACHE(LFC hits/misses) as the load-bearing observability layer for diagnosing storage-compute-boundary performance issues. New patterns (2): patterns/snapshot-plus-catchup-replication (the 10% / 10× rule of thumb generalised — choose snapshot mode when the per-cycle delta exceeds the implementation-specific crossover threshold, where bulk-copy efficiency overtakes per-row diff/merge cost; counterintuitive because snapshot rewrites the entire table, but for high-delta workloads the bulk-copy path wins by an order of magnitude on the disclosed Lakebase implementation); native-postgres-roles-for-non-databricks-aware-partners (OAuth's hourly token rotation incompatibility with partner systems like SAP Engagement Cloud forces a fall-back to native Postgres password roles with operator-managed rotation; explicit security-discipline-instead-of-mechanism tradeoff). Lakebase page extension: seventh canonical face for Lakebase added to systems/lakebase — the bidirectional governed-data plane between operational and analytical tiers, with LFC as the compute-side observability layer for the storage-compute boundary. Concrete sizing disclosure: bursty marketing-campaign workload at Deichmann uses scale-to-0 → 16 CU (~32 GB RAM) on Lakebase Autoscaling, with the architectural justification that "Lakebase autoscaling speed and reactivity eliminate the risk of resource underutilization" — sub-second scale-down makes generous max-cap sizing safe. Marketing-campaign customer segments positioned as the canonical bursty workload applied to OLTP rather than to observability databases (the prior canonical context). Standard Postgres tuning surface (pg_stat_statements,work_mem256 MB on larger compute,autovacuum_vacuum_scale_factorfor high-churn tables) confirmed unchanged. TLS chain: Let's Encrypt; partner systems must trust ISRG Root X1. -
2026-05-20 — sources/2026-05-20-databricks-governing-ai-agents-at-scale-with-unity-catalog (Tier-3 Databricks blog vision/positioning post extending the 2026-04-17 Unity AI Gateway launch from coding-agent scope to org-wide agent governance across every department — dev / analytics / sales-ops / support / marketing / finance). Borderline-include: vision-heavy with no scale numbers, no internals, but the four-pillar framing and five named architectural extensions (Service Policies, Inference Tables, Lakewatch, Guardrails, Budgets) are individually citable. Canonical four-pillar framing for agent governance ((1) Delegated access — three-layer permissions/Service-Policies/Guardrails composition; (2) Data-centric AI governance — "AI governance is data governance"; (3) Cost intelligence — usage-tracking + Budgets; (4) Open and interoperable — governance-travels-with-resources). New systems (5): Service Policies (UC functions attached to registered MCPs evaluating tool calls before execution; ternary
allow/deny/consent; fail-closed on deny; canonical instance of policy-as-uc-function-attached-to-mcp); Inference Tables (full payload of every model call — "the exact prompt sent, the exact response returned, token counts and latency" — written to UC-managed Delta tables, customer-controlled retention; canonical instance of inference-payload-table-for-audit); Lakewatch (Databricks' agentic SIEM built on the security lakehouse — first wiki disclosure; "Attackers are using agents. Defenders should too."); Unity AI Gateway Guardrails (inline content scanning of every model call — inputs for PII + jailbreak, outputs for hallucinations + sensitive content; fail-closed on every request; canonical instance of concepts/ai-agent-guardrails); Unity AI Gateway Budgets (per-user / per-group monthly spend thresholds with alerts; hard enforcement on roadmap). New concepts (4): concepts/centralized-ai-governance (the canonical four-pillar framing distinct from the 2026-04-17 three-pillar shape — adds delegated-access and open-interoperable as load- bearing additions); concepts/centralized-ai-governance ("an agent's behavior is almost entirely determined by the data it has access to" → AI governance and data governance must be one system; the data-classification → tag → ABAC pipeline applies to agent traffic without AI-specific configuration); concepts/ai-agent-guardrails (canonical concept distinct from CI/quality-gate-based concepts/ai-agent-guardrails; bidirectional, inline, per-request, fail-closed scanning of LLM I/O); governance-travels-with-resources (Pillar 4 principle: "governance becomes a property of your platform rather than something you rebuild for each new framework or model" — same UC + AI Gateway across LangGraph / CrewAI / OpenAI SDK / Anthropic SDK / AutoGen / LlamaIndex; same gateway across Databricks-hosted / Azure OpenAI / AWS Bedrock / Anthropic). New patterns (3): three-layer-agent-control (the load-bearing composition for Pillar 1 — permissions key on identity, Service Policies key on tool-name+args+identity, Guardrails key on payload content; each layer's decision input is strictly weaker than what the next layer needs); inference-payload-table-for-audit (full request/response capture in lakehouse-resident tables — breaks the conventional completeness-vs-cost tradeoff that APM-style logging imposes); policy-as-uc-function-attached-to-mcp (catalog-managed policy code attached to the resource (MCP server), not the agent — so framework-agnostic; UC functions inherit UC's versioning + audit + ownership lifecycle). Extends 7 existing pages: systems/unity-catalog (eleventh face — AI-asset-governance substrate for "LLMs, MCP servers, skills, and agents"; first explicit end-to-end OBO disclosure — "identity flows… from the user who asks the question to the specific table row the agent retrieves" — with dual-identity audit logging); systems/unity-ai-gateway (org-wide-agent generalisation + five new architectural surfaces: Service Policies, Guardrails, Inference Tables, Budgets, Lakewatch substrate); systems/model-context-protocol (MCP servers now registered in UC as securables, governed with Service Policies + OBO); concepts/centralized-ai-governance (three-pillar → four-pillar comparison table); concepts/governed-agent-data-access (Databricks instance of Gallego's two-axis framing, with three-layer + four-pillar framings as more-granular siblings); coding-agent-sprawl (generalised to org-wide agent sprawl across every department; talent-flight as a third risk axis); patterns/on-behalf-of-agent-authorization (Databricks generalisation + "specific table row" granularity disclosure + dual-identity audit logging requirement); patterns/central-proxy-choke-point (generalised from coding-agent scope to org-wide agent scope; canonicalises the three-layer-agent-control composition the choke-point hosts); patterns/telemetry-to-lakehouse (full-payload specialisation via Inference Tables — distinct from but co-existing with the metrics+traces variant). - 2026-05-20 — sources/2026-05-20-databricks-virtue-foundation-medical-volunteers-72-countries (Tier-3 Databricks-for-Good co-marketing post — Virtue Foundation's VF Match platform connects medical volunteers to opportunities in 72 low / low-middle income countries via a production-grade Foundational Data Refresh (FDR) pipeline on Databricks). Borderline-include: ~40% Databricks-for-Good framing in the bookend paragraphs but the Building the Foundation / Entity Resolution at Scale / VF Agent sections (~60%) name a specific architectural shape worth canonicalising — a multi-step LLM extraction pipeline over 25M+ web pages orchestrated by Lakeflow Jobs across 15+ interdependent tasks, with star-schema state + status-based checkpointing + a configurable extraction registry; an Entity Resolution stage built on the open-source Splink probabilistic record-linkage framework with a quantified curse-of-the-last-reducer observation (one Spark partition running 30 minutes vs 52-second median — ~35× ratio) reduced 15× to ~2 minutes by enabling Photon; and a prototype VF Agent multi-agent architecture in LangGraph routing user queries to Vector Search or Genie sub-agents via a Multi-Agent Supervisor. New systems (6): systems/vf-match (Virtue Foundation's volunteer-matching marketplace, the user-facing system); systems/vf-agent (prototype natural-language-query layer with four-sub-agent composition on LangGraph); systems/splink (UK-Ministry-of- Justice-origin Apache-2.0 probabilistic record-linkage framework implementing Fellegi-Sunter + EM-driven match-weight estimation, SQL-pluggable backend across Spark / DuckDB / others); first dedicated wiki page for systems/photon (Databricks' C++-native vectorised query engine, previously only mentioned in passing across multiple ingests); systems/langgraph (LangChain ecosystem's graph-based agent-orchestration framework — first wiki disclosure); systems/overture-maps + systems/bright-data (the two complementary data sources feeding FDR — Meta+Microsoft open-source geospatial authority
- commercial real-time web scraping). New concepts (6): multi-step-llm-extraction (the break-the-task-into- targeted-steps discipline at LLM-invocation altitude; "dramatically reduces token consumption while focusing each model invocation on a narrow, high-precision task"); status-based-llm-pipeline-checkpointing (per-record state-column primitive enabling 25M+-page pipeline resumability without re-paying LLM cost); concepts/oltp-vs-olap (the fact-and-dimension data-warehousing schema as state model for LLM pipelines — first wiki canonicalisation of star-schema as the substrate for multi-step LLM extraction state); probabilistic-record-linkage (Fellegi-Sunter formal-statistical formulation Splink implements); vectorized-query-engine (the engine class Photon / DuckDB / Velox / ClickHouse all implement, with batch-of-rows SIMD column-major loops); concepts/partition-skew-data-skew (the canonical name for straggler-partition-dominates-wall-clock at the reduce stage of batch jobs). New patterns (2): multi-step-llm-extraction-pipeline (named pattern composing the multi-step + status-checkpointing + registry + star-schema sub-properties; first canonical wiki instance of the production-grade LLM-extraction pipeline shape distinct from the document-extraction sibling patterns/two-stage-evaluation); patterns/specialized-agent-decomposition (one supervisor classifies query intent + complexity and routes to one of N alternative-answer-shape sub-agents; distinct from Claroty's collaborative role-decomposition which has agents collaborate on one task — supervisor-routing has agents as alternatives selected per query). Updated (8 existing pages): concepts/entity-resolution gains Splink as the canonical open-source classical-ER framework (first open-source-ER instance on the wiki; complementary to Claroty's custom-built hybrid stack); concepts/partition-skew-data-skew gains the Spark / Photon batch-altitude instance with the 30 min → 2 min remediation (15× improvement quantification via vectorisation rather than partition redistribution); systems/lakeflow-jobs gains the third canonical face (FDR 15+-task multi-step LLM-extraction-pipeline orchestrator alongside MapAid groundwater + Claroty CSAF — the shape is converging across three independent customers); systems/databricks-genie gains the fifth canonical face (Genie-Agent-as-sub-agent) — Genie used not as a destination chatroom but as an internal subroutine in a larger query-orchestration graph, distinct from BI-replacement / data-agent-internals / embedded-NL-query / migration-handoff faces; systems/mosaic-ai-vector-search gains the Vector-Search-Agent-in-multi-agent-supervisor face; systems/apache-spark gains the LLM-extraction-pipeline + ER face (25M+-page extraction + Splink ER, with the curse-of-the-last-reducer observation pinning Photon's value at non-OLAP altitude — first wiki-quantified instance); companies/databricks tags + Recent articles extended. Operational numbers cited: 72 countries; 50,000+ patients delivered care to date; 25M+ web pages processed through GPT models; 15+ interdependent Lakeflow Jobs tasks; 30 min worst- case partition vs 52 s median (~35× ratio) before Photon; ~2 min worst-case partition after Photon; 15× improvement; 4 named sub-agents in VF Agent prototype. Architectural composition summary: the same Databricks-stack primitives that compose into MapAid groundwater (multi-step LLM extraction on scanned documents) and Claroty CPS (hybrid ER on industrial asset identity) compose into VF Match (multi-step LLM extraction on web pages + Splink ER) — the **multi-step-LLM-extraction
-
entity-resolution + multi-agent-query-layer** trifecta is the recurring architectural shape on Databricks for AI-powered data catalog construction. Caveats: vendor co-marketing post; VF Agent is a prototype; per-step accuracy / cost / latency numbers not disclosed; GPT model versions / per-step prompts not disclosed; only the Photon comparison is quantified.
-
2026-05-19 — sources/2026-05-19-databricks-deutsche-borse-zeppelin-to-databricks-notebook-migration (Tier-3 Databricks Blog, customer-co-authored with Deutsche Börse). Borderline-include launch-post-with-architecture: ~30% architecture density passes the AGENTS.md "Borderline cases — include, don't skip: Product launches THAT ALSO contain deep architecture sections" test, with named-primitive disclosure (paragraph-to-cell + interpreter-prefix mapping +
.ipynbJSON reformat as the deterministic stage; Genie + context-encoded prompt as the LLM stage; explicit negative space — "the converter does not rewrite SQL logic, Python logic, visualizations, widgets, Oracle and HDFS references, scheduling logic or business- specific custom code") and a concrete reusable architectural thesis: "separate structure from logic, apply the right tool to each." Forcing function: Cloudera Zeppelin EOL 2027. Created (8 new pages): 1 source + 1 company (companies/deutsche-borse) + 2 systems (systems/apache-zeppelin, systems/deutsche-borse-zeppelin-converter) + 3 concepts (notebook-format-migration, heterogeneous-code-migration, concepts/context-engineering) + 2 patterns (structural-deterministic-logical-llm-split, patterns/context-segregated-sub-agents). Extended (2 systems): systems/databricks-apps gains a fourth canonical face — customer-built migration-tool substrate — distinct from clinical-ops decision-support / Claroty HITL-UI / DBA- automation; the App is itself a migration utility, not a destination workload, and runs inside the destination platform's workspace. systems/databricks-genie gains a fourth canonical face — migration-handoff: consumer of a context-encoded prompt emitted by a deterministic operator-side tool — distinct from BI-replacement / data-agent-internals / embedded-NL-query; pins Genie's effectiveness to upstream context-engineering discipline for a third time on the wiki (alongside Trinity measure- consolidation as load-bearing precondition and the 2026-05-08 rich semantic enterprise context as substrate). Operational results: hours-to-minutes per notebook (manual: hours each → hybrid: 15–20 min each), 2,000-user migration scope, business- user-self-service workflow with no per-notebook engineering team. Lessons-learned section also pins a notable counter- cyclical signal: a first-attempt agentic architecture was explicitly rejected in favour of a simple linear two-stage pipeline (Stage 1 → handoff → Stage 2), on the rationale that "a more complex agentic architecture added overhead without solving the core problem". Frontend stack disclosure: shadcn UI in production, evolved from a Streamlit prototype. -
2026-05-15 — sources/2026-05-15-databricks-backstage-with-lakebase-part-2 (Tier-3 Databricks Blog, Part 2 of the Thoughtworks Backstage- with-Lakebase series — Governance). Borderline-include: ~70% architecture density passes the AGENTS.md test, with named- primitive disclosure across Lakehouse Federation (operational Postgres exposed as foreign catalog
lakebase_bsin UC),system.access.audit(every Lakebase control-plane action), UC system billing tables (cost attribution by(project_id, branch_id, endpoint_id)with worked numbers31.6130 DBUprod /0.0107 DBUtransient test branch), branch-propagated masking policies (UC attribute-level masks inherit at branch creation), and two open-source Thoughtworks tools deployed as Databricks Apps: LakebaseOps (three- agent platform — Provisioning / Performance / Health — replacing 51 historical DBA tickets, 7 scheduled Databricks Jobs replacing pg_cron, 9-KPI adoption dashboard, ten-engine migration wizard with live AWS+Azure pricing) and Lakebase MCP (46-tool MCP server with dual-layer governance — SQL-statement guard + per-tool access guard across four profilesread_only/analyst/developer/adminmapping onto UC GRANT — plus per-statement tool-tag attribution). Load-bearing claim: "Because Lakebase is natively embedded inside Databricks, Unity Catalog extends directly over the operational Postgres database." The compliance side-channel (CloudTrail + pgaudit + CloudWatch) collapses into one SQL query against UC system tables. Created (10 new pages): 1 source + 3 systems (systems/lakebaseops, systems/lakebase-mcp, systems/lakehouse-federation) + 3 concepts (branch-level-cost-attribution, concepts/centralized-ai-governance, operational-analytical-governance-unification) + 3 patterns (foreign-catalog-federation-for-operational-db-governance, dual-layer-governance-sql-and-tool-guards, tool-tagged-query-attribution). Extended (5 pages): systems/lakebase (8th face — governance- substrate-unified operational DB), systems/unity-catalog (10th face — operational-DB governance via Lakehouse Federation), systems/backstage (2nd face — governance- substrate-unified IDP), systems/databricks-apps (3rd face — DBA-automation + AI-agent-DB-access deployment substrate), concepts/database-branching (governance-composition leg of Part 2). Eighth Lakebase face on the wiki. Cross-source continuity: direct sequel to Part 1 (Deployment Cycles) which canonicalised branching as developer-cycle primitive; Part 2 makes governance inseparable from branching. Companion to the 2026-05-13 UC ABAC GA — same UC ABAC primitives apply to the federated operational DB. Part 3 (FinOps) is forthcoming and explicitly previews "taking the infrastructure ownership data inside Backstage and joining it directly to cloud billing data in a single SQL query." -
2026-05-14 — sources/2026-05-14-databricks-expanded-interoperability-with-unity-catalog-open-apis (Tier-3 Databricks Blog launch post — Expanded interoperability with Unity Catalog Open APIs). Borderline-include: launch framing but ~40% architecture density passes the AGENTS.md borderline test, with named-substrate disclosure across catalog-managed commits, credential vending, and Delta Kernel. Two coordinated milestones: External Access to Managed Tables in Beta (Apache Spark / Apache Flink / DuckDB can create, read, write, and stream to/from UC managed Delta tables) and Credential Vending GA for tables / Public Preview for Volumes (M2M OAuth replaces PATs; engine-side auto-refresh closes the long-running-pipeline gap). Three new system pages (systems/uc-managed-tables — the open-API-accessible managed Delta table primitive with Predictive Optimization + Liquid Clustering + catalog commits; systems/uc-credential-vending — the credential-vending API with M2M OAuth + auto-refresh; systems/delta-kernel — the open-source Java + Rust library abstracting the Delta protocol so engines focus on UC integration; plus first wiki disclosure of systems/duckdb as a UC-integrated external engine). Four new concept pages (concepts/short-lived-credential-auth — first wiki canonicalisation of the catalog-mediated short-lived-scoped-credential primitive; concepts/open-table-format — first wiki canonicalisation of the central commit-coordinator substrate that prevents log corruption + provides audit + enables multi-table transactions; concepts/open-table-format — first wiki canonicalisation of the architectural shape that resolves the managed-table-benefits + compute-engine-choice trade; m2m-oauth-vs-pat — first wiki canonicalisation of the auth-substrate comparison naming the three structural failure modes of PATs that M2M OAuth dissolves). Three new pattern pages (credential-vending-for-external-engine-access — auth-side deployment pattern with the two-layer short-lived- credential shape; catalog-managed-commits-for-external-write-safety — commit-side deployment pattern naming serialized commits + audit + multi-table transaction substrate as the three property guarantees; connector-library-as-protocol-abstraction — first wiki canonicalisation of the library-shape pattern with Delta Kernel as canonical instance, naming the two structural failure modes that protocol-abstraction libraries eliminate (drift across implementations + per-engine reimplementation cost)). Updated five existing pages: systems/unity-catalog gains a ninth canonical face (Open API + external-engine-write hub) with full architectural-primitives disclosure; systems/delta-lake gains a ninth canonical face (external- engine-managed-write substrate) — managed Delta tables now writeable by Spark / Flink / DuckDB via Delta Kernel under UC catalog commits; systems/apache-spark + systems/apache-flink gain UC-Managed-Table-writer faces; systems/unity-catalog-volumes composes via Volume Credential Vending (Public Preview). Forward-roadmap composition with ABAC for external reads: the post explicitly names ABAC for external reads as developed-but-not-yet-GA, which when shipped will compose the 2026-05-13 ABAC GA primitives with the 2026-05-14 external-access primitives — fine-grained row + column-level governance enforced uniformly across first-party and external read paths. PepsiCo testimonial (Sudipta Das, Director of Enterprise Data Operations) frames the customer payoff. Activation contract: preview-portal enrollment + metastore-level toggle + schema-level
EXTERNAL_USE_SCHEMAgrant + Delta-Spark 4.2 / UC 0.4.1 version pinning. -
2026-05-13 — sources/2026-05-13-databricks-the-rosetta-stone-of-cps-clarotys-ai-powered-library (Tier-3 Databricks Blog co-marketing post — Databricks GenAI MVP customer story on Claroty's AI-Powered CPS Library, the asset-identity layer behind Claroty's xDome CPS-protection platform). Borderline-include: heavy GenAI-MVP framing (~70%) but the Under the Hood / Data Engineering at Scale / Multi-Agent Intelligence / Innovation through Databricks Capabilities sections (~30%) name a specific architectural shape worth canonicalising — a hybrid Entity Resolution pipeline combining classical ER with an orchestrated multi-agent system (NLP / Reasoning / Human-in-the-loop), all on the Medallion Architecture over Delta Lake with Delta Change Data Feed driving a dynamic mapping registry, plus a real production observation about vector-search endpoints lacking scale-to-zero for bursty workloads. One new system page (systems/claroty-cps-library — Claroty's xDome asset-identity layer at 17 million+ asset catalog scale; CPS-ID positioned as "the new industry standard for cyber-physical system identity"). Three new concept pages (concepts/entity-resolution — first wiki canonicalisation of ER as an architectural problem class with the noisy-real- world-data-into-single-source-of-truth shape; concepts/change-data-capture — first wiki canonicalisation of Delta CDF as the layer-transition trigger driving Bronze → Silver promotion via dynamic mapping-registry application; concepts/scale-to-zero — first wiki canonicalisation of the explicit production-cost observation that hosted vector-search endpoints currently lack scale-to- zero, structurally analogous to concepts/scale-to-zero but at the index altitude). Two new pattern pages (hybrid-classical-er-plus-genai — the central architectural shape combining battle-tested classic ER with GenAI cognitive parsing as complementary halves; orchestrated-multi-agent-entity-resolution — the three-role decomposition NLP / Reasoning / HITL with the feedback loop closing into model retraining). Five existing systems extended: systems/delta-lake gains a Delta-CDF
- schema-evolution + time-travel audit-chain face;
systems/lakebase gains the transactional ER store face
(sixth-or-later wiki face — Postgres constraints load-bearing
for asset-mapping data integrity); systems/databricks-apps
gains the HITL UI for Entity Resolution face composed
with Lakebase as state store; systems/databricks-model-serving
gains the substrate for hosting domain-specific embedding
models as custom endpoints face (third wiki face — beyond
the 200K-QPS Superhuman LLM-serving face); systems/mlflow
gains the continuous-production-monitoring substrate against
concept drift via LLM-as-a-Judge face; systems/lakeflow-jobs
gains the CSAF security-advisory ETL with
ai_query+ step-by-step Delta tables face; systems/databricks-ai-functions gains the second canonical instance (Claroty CSAF pipeline alongside MapAid groundwater pipeline). Three existing concept pages extended: concepts/medallion-architecture gains a third canonical instance (Bronze raw → mapping- registry-canonicalised silver schema for ER); concepts/llm-as-judge gains a third face (production- monitoring against concept drift in CPS data with conservative pass/fail/unknown ternary explicit against missing ground- truth); concepts/schema-evolution gains the audit-chain- enabler-with-time-travel face. One existing system extended: systems/unity-catalog gains an eighth wiki face — Entity-Resolution catalog governance, the audit-chain anchor connecting raw evidence → mapping-registry version → canonical CPS-ID → vulnerability attribution. Operational numbers cited: 17M+ assets in the global catalog; 88% of CPS assets do not transmit an exact product code; 76% transmit codes that differ from vendor records; 25% improvement in vulnerability identification accuracy; 56% of analysed devices receive new or updated security recommendations for previously-invisible outdated firmware; worked example "Rockwell Automation 1769-L36ERMS/B → Compact GuardLogix 5370 → CVE-2020-6998". Architectural composition summary: the same Lakebase + UC + Apps + Model Serving + MLflow primitives that compose into the clinical-operations decision-support shape (FPW / 2026-05-13 clinical-ops source) compose differently into the Entity Resolution catalog shape here — same primitives, different role-assignment (Lakebase is ER asset store not app shortlist store; Apps is SME HITL UI not clinician decision- support UI; Model Serving hosts custom medical/OT embedding endpoints not LLMs at 200K QPS). The recurring shape: a small set of Databricks platform primitives composing into very different vertical workflows by re-assigning their roles. - 2026-05-13 — sources/2026-05-13-databricks-clinical-operations-intelligence-belongs-on-the-lakehouse (Tier-3 Databricks Blog — open-source release of the Site Feasibility Workbench as a Databricks App for clinical-trial site selection, framed as a reference implementation of "clinical operations intelligence when the application, the models, and the data live on the same platform."). Borderline- include: ~30% architecture density (the Architecture Argument
- Auditability Argument sections name specific platform primitives and articulate a concrete architectural thesis), ~70% clinical-trial industry context. Two new system pages (systems/databricks-apps — first wiki disclosure of Apps as a deployment model; systems/site-feasibility-workbench — open-source reference implementation, FastAPI + React, ~30 min deployment) plus two new concept pages (single-platform-application-architecture — the unified-platform thesis that eliminates four integration layers (sync pipeline + credential surface + RBAC translation + semantic harmonisation); governed-shap-attribution-table — per-prediction SHAP attributions stored as governed UC Delta tables, versioned in MLflow, lineaged through UC, queryable in SQL) plus two new pattern pages (in-workspace-app-as-decision-support — the deployment shape; shap-attribution-as-governed-delta-table — the regulated-ML audit pattern). Architectural thesis: "Databricks Apps, Lakebase, and AI/BI Genie eliminate each of those layers — not by abstracting them away but by making them unnecessary." The single-platform shape is a four-primitive composition: Apps (web tier inside the workspace, service- principal auth, SQL Statement API path to UC, REST API to Genie, all internal connections); Lakebase (operational app state, scale-to-zero, workspace-credentialed); UC (governance + lineage
- RBAC the app inherits for free, plus the substrate for governed SHAP attributions); Genie (embedded NL-query layer). Sixth Lakebase face on the wiki (after CMK / LangGuard agentic-OLTP / Stripe-Projects agent-provisioning / Backstage state-heavy-app / FPW image-generation-pushdown): app-tier-state-store-without-its- own-credential-surface. Seventh UC face: data-plane half of the single-platform application architecture. New AI/BI Genie face: embedded NL-query layer composed via internal REST API, not a separate Genie-room product. New MLflow face: the versioning leg of the regulated-ML SHAP-attribution-Delta-table audit substrate (anchors per-prediction attributions to the exact model version, addressing 21 CFR Part 11 + ICH E6(R3) + FDA GMLP audit-chain integrity). New Delta Lake face: ML-audit substrate for governed per-prediction SHAP attributions with the three load-bearing properties (ACID + schema evolution + SQL queryability + time-travel). Architectural reframe of explainability as fairness control: per the source, "sponsors can audit recommendations for systematic under-weighting of community sites, minority-serving institutions, or first-time investigators — turning explainability into a fairness control" — the substrate enabler is the queryable-attribution population, not the per-prediction explainer service. Reference implementation deploys "into an existing Databricks workspace with Unity Catalog in approximately 30 minutes of technical deployment time." Forward roadmap (named in post): three additional Databricks Apps (Patient Cohort and Recruitment, Enrollment Velocity Optimizer, Risk-Based Monitoring and Compliance) on the same shape — "All four deploy as Databricks Apps. All four query Unity Catalog directly. None make external API calls."
- 2026-05-13 — sources/2026-05-13-databricks-abac-row-filtering-and-column-masking-policies-governed-tags
(Tier-3 Databricks Blog — GA announcement for three Unity
Catalog governance primitives co-designed into one
organize → detect → protect pipeline). Three new system pages
(UC ABAC, UC
Governed Tags, UC
Data Classification) plus four new concept pages
(governed-tag, data-classification-tagging,
concepts/least-privileged-access,
session-identity-evaluation) and two new pattern pages
(tag-driven-attribute-based-access-control,
single-variant-udf-for-multi-type-masking)
canonicalise the GA primitives. Sixth wiki face for Unity
Catalog: not just where governance is recorded but where it
is expressed, evaluated, and enforced. Architectural shift is
from per-object configuration of row filters / column masks
("repetitive and prone to inconsistency") to declarative
tag-driven policy evaluation — one ABAC policy referencing
pii:ssncovers every column carrying that tag in scope, including columns added or tagged after policy authoring. Three GA scaling enhancements: (1) policy limits grew 10× across every scope (10K+ per metastore, 100+ per catalog/schema); (2) session identity evaluation for views and functions evaluates against the user running the query, not the view creator (closes view-as-bypass failure mode); (3) single VARIANT UDF can maskINT/DOUBLE/DECIMAL/STRUCTcolumns at once via type-erasure (single-variant-udf-for-multi-type-masking). Built-in data classifiers cover GDPR / HIPAA / GLBA / DPDPA / PCI plus UK / Germany / Australia / Brazil regional packs (India + Canada coming May 2026). Custom classifiers in Beta learn detection patterns from already-tagged columns + Unity Catalog metadata, with human-in-the-loop false-positive exclusion feeding back to improve precision. Separation-of-duties (concepts/least-privileged-access) emerges from the three-permission split across the primitives: governance teams hold tag-taxonomyMANAGE/CREATE, stewards hold tagAPPLY, data producers hold tableOWNER— three roles operating on three permission axes without cross-team blocking. Architectural composition with prior wiki framings: (a) UC ABAC is ABAC at the table-storage governance altitude, distinct from the prior-canonical Convera + Cedar API-authorization altitude; (b) UC Governed Tags is the data-warehouse-catalog altitude of data-classification-tagging, distinct from Figma FigTag's application-schema altitude and Meta Policy Zones' runtime-IFC-annotation altitude; (c) the organize → detect → protect pipeline mechanically fills in the "data classification with governed tags + row/column-level controls" governance contract that the multimodal-healthcare ingest named for governed-delta-tables-per-modality; (d) broadens the principal surface Genie operates over (the GA post explicitly cites Genie as a driver — agents make per-table per-user wiring even less tractable, motivating tag-driven policy). Borderline Tier-3 (GA announcement, ~70% architectural / ~30% PR) — included because the governance-architecture content (organize → detect → protect with no handoff, ABAC over governed tags, separation of duties, session identity, VARIANT UDF type-erasure) is substantive even though implementation depth is light. Customer testimonials: Atlassian (Gerald Nakhle, operational-overhead reduction), Udemy (Rajit Saha, "Fewer policies, lower costs, surgical precision"), Superhuman (Nan Wu, on agentic Data Classification "replaces manual overhead with automated, high-quality results"). - 2026-05-11 — sources/2026-05-11-databricks-unlocking-the-archives
(Tier-3 Databricks Engineering / Databricks-for-Good post —
first wiki disclosure of
AI Functions (
ai_query) used as the universal inference primitive across three pipeline stages, plus first instance of Databricks Asset Bundles and Foundation Model API on the wiki). MapAid + SUDAAK groundwater archive: ~700 scanned PDFs / 5,570 pages of decades-old, mixed English/Arabic Sudanese geological surveys turned into a structured catalog + 299 well/borehole records via a multimodal classification + judge + extraction pipeline. Three architectural primitives disclosed: (1) Visual-First Document Extraction — page-as- image to multimodalai_query, OCR replaced by visual understanding (concepts/multi-modal-attribute-extraction + visual-first-document-extraction). (2) Two-Pass Classify-then-Deep-Extract with Intelligent Sampling — sample title pages / intros / conclusions in pass 1 for >70% volume reduction; full-page OCR + entity-anchored linking + JSON extraction only on the ~50% water-flagged subset in pass 2 (patterns/two-stage-evaluation). (3) LLM-as-Judge as a First-Class Pipeline Stage — every classification scored on accuracy / completeness / consistency rubric with categorical rating + written justification (audit trail), sub-threshold rows routed to manual review (llm-judge-as-inline-pipeline-stage). Also canonicalises schema- constrained LLM output, [[patterns/sql-native-multimodal-llm- inference|SQL-native multimodal inference]], and Asset Bundle single-command deployment. First-run numbers: 654 docs / 5,570 pages in <3 hours, 95% rated excellent/good by inline judge, 299 structured records extracted. Borderline Tier-3 (Databricks-for-Good customer story) — included because the architectural primitives generalise. Eighth Databricks-platform face on the wiki. -
2026-05-08 — sources/2026-05-08-databricks-pushing-the-frontier-for-data-agents-with-genie (Tier-3 Databricks Engineering post — first mechanism-level disclosure of Databricks Genie's internal architecture beyond the product-name level, building on the 2026-04-29 Trinity Industries adoption case). Three named architectural advances: (1) Specialised Knowledge Search — Genie "uses the existing data assets such as workspace tables, notebooks, dashboards, documents, and files to derive a rich semantic enterprise context and then uses this context to construct a search index. It uses multiple search indices in parallel together with rich metadata signals." Disclosed result: "up to 40% improvement on table discovery benchmarks" (Figure 4). Canonicalised as specialized-knowledge-search + semantic-context-grounded-search-index over the rich semantic enterprise context substrate. (2) Parallel Thinking — sample multiple agent trajectories over the same query, aggregate findings; "significant accuracy improvement" (Figure 5) on GPT-5.4 + Opus-4.6 baselines. Canonicalised as parallel-thinking-trajectory-sampling + parallel-trajectory-sampling-and-aggregation, positioned as the structural compensation for the verifiable-test gap unique to data agents. (3) Multi-LLM — different LLMs per sub-agent (planning / search / code-gen / judges) with GEPA-optimised prompts per (LLM, sub-task) pair; combined effect "significantly reduce costs and latency" simultaneously with the accuracy gain. Canonicalised as multi-llm-sub-agent-routing + llm-per-subagent-with-optimized-prompts. Four- phase trajectory disclosed via worked example (CFO question about contradictory revenue dashboards): (1) parallel multi- agent data discovery, (2) data investigation (SQL extraction + comparative + root-cause), (3) self-correction loop / reconciliation (concepts/agentic-development-loop), (4) verification. Canonicalised as four-phase-data-agent-trajectory. Headline operational result: Genie 32% → over 90% accuracy vs "a leading coding agent" (name not disclosed) on Databricks' internal benchmark of real-world data-analysis tasks — claimed simultaneously on accuracy + cost + latency. First wiki naming of the data-agent vs coding-agent distinction (data-agent-unique-challenges) with three structural challenges: scale of data discovery, source-of-truth disambiguation (concepts/data-lineage), and the verifiable-test gap. The post also makes the load-bearing dependency on upstream governance discipline mechanically precise — Genie derives its semantic context from existing workspace assets, so the prior Trinity-Industries empirical disclosure (effectiveness depends on prior measure-consolidation work) is now an architectural property: if the data layer hasn't disambiguated, Genie's specialised search has nothing rich to ground on. Architectural-canonicalisation contributions: (1) first mechanism-level Genie internals disclosure beyond Trinity adoption framing; (2) first wiki naming of data agent as a class structurally distinct from coding agent, with three uniquely-data-agent challenges; (3) first wiki canonicalisation of parallel thinking via trajectory sampling + aggregation as a named agent-design technique compensating for missing oracles; (4) first wiki canonicalisation of Multi-LLM sub-agent routing as a per-sub-agent assignment shape (vs prior model-cascade / routing-layer framings); (5) first wiki canonicalisation of GEPA (systems/gepa-prompt-optimizer) as a referenced production prompt-optimisation tool; (6) first wiki canonicalisation of the four-phase data-agent trajectory shape distinct from coding-agent write-test-iterate loops; (7) first wiki canonicalisation of the self-correction loop as a load-bearing intra-trajectory mechanism; (8) first wiki canonicalisation of the "agent-architecture choices recover all three of accuracy + cost + latency simultaneously" Pareto move (counter to the typical assumption that sampling trades cost for accuracy). Caveats: specific (LLM, sub-agent) assignments not disclosed; trajectory count + aggregation strategy not disclosed; self-correction mechanism not disclosed; internal-benchmark composition not disclosed; "leading coding agent" baseline name not disclosed; no latency / QPS / cost numbers for Genie endpoints. Sibling to the prior 2026-04-29 Trinity Industries adoption ingest (Genie deployed empirically) and the 2026-04-29 Stripe Projects + Databricks launch (Lakebase as substrate for agents) and the 2026-05-05 / 2026-05-06 / 2026-05-07 / 2026-05-08 Databricks platform-engineering arc (Pantheon-Hydra observability + Spark-Connect serverless OLAP + Lakebase OLTP + Model-Serving managed inference). Together the five-article 2026-05 Databricks engineering window now spans observability + OLAP + OLTP + managed-inference + data-agent altitudes in nine days.
-
2026-05-08 — sources/2026-05-08-databricks-how-superhuman-and-databricks-built-a-200k-qps-inference-platform-together (Tier-3 Databricks Engineering post — first canonical wiki disclosure of Databricks Model Serving internals at the platform-engineering altitude; Databricks' fourth architectural-retrospective in nine days, extending the 2026-05 Databricks engineering ingest window across the OLTP (Lakebase) + OLAP (Serverless Compute) + observability (Pantheon/Hydra) tiers and now into managed external inference). Joint engineering retrospective with Superhuman documenting the migration of the grammar-correction LLM (40M+ daily users, peak 200,000+ QPS, sub-1s p99, 4-9's reliability, ~50/50-token request shape) off a self-managed vLLM-on-L40S DIY stack onto Databricks Model Serving on H100. Two-layer co-engineering: Platform layer — EDS-driven P2C load balancer (replacing default Kubernetes round-robin which "degrades at higher QPS"),
request_concurrency-based asymmetric autoscaler (aggressive scale-up, conservative scale-down for anti-flapping), block-device lazy-loading container image with 4MB sectors cutting pod start from "several minutes to a few seconds". Runtime layer — FP8 quantisation (single largest win, up to +30% per-pod QPS) with attention (Q/K/V/output) + MLP projections on FP8 path, KV-cache quantisation explicitly disabled ("weight quantization was where the throughput wins came from and KV-cache quantization introduced its own quality tradeoffs that weren't worth pursuing for this workload"); ** per-channel FP8 scaling beating off-the-shelf per-tensor scaling at matched throughput; hybrid- precision toggle so attention quantisation can be flipped on/off via flag without architectural change; multiprocessing RPC server (+20%) addressing the CPU-bound regime small fast LLMs hit on H100; single-call C++ tensor manipulation in the CUDA-graph decode step + an async CPU-GPU scheduler overlapping batch-N post-processing with batch-N+1 forward pass. Net per-pod throughput: 750 → 1,200 QPS on H100 (+60%). Architectural-canonicalisation contributions: (1) first wiki disclosure of Databricks Model Serving's internals beyond the product-name level; (2) first wiki disclosure of the CPU-bound regime for small fast LLMs with a named workload (Superhuman grammar correction); (3) first wiki canonicalisation of the managed-serving-without-giving-up- control division of responsibilities (customer owns model + quantisation + quality bar; platform owns runtime + LB + autoscaler + image substrate); (4) first wiki disclosure of KV-cache-quantisation-explicitly-off as the production-LLM- serving selective-FP8 boundary; (5) first wiki canonicalisation of the autoscaler-altitude anti-flapping primitive (asymmetric scale-up/scale-down to prevent the latency-spike flapping cycle); (6) first wiki canonicalisation of container-image cold start* as a fourth serverless cold-start regime distinct from CPU-runtime-init / V8-isolate / GPU-weights- load. Joint shadow testing was the joint-engineering primitive used to tune autoscaler thresholds and validate quality on Superhuman's internal eval harness with zero quality regression. vLLM stayed in the toolchain post-migration as the prequantisation library (Superhuman's ML team prequantised the FP8 checkpoint "using vLLM's online quantization library"*). Caveats noted: 200K-QPS / sub-1s-p99 / 4-9's claims are about the Superhuman endpoint specifically (not Databricks Model Serving in general); per-pod 1,200 QPS is at the 50/50-token shape on H100; quality validation is on Superhuman's internal eval harness (not a public benchmark); the L40S baseline was not separately re-tested with the new optimisations so some of the throughput improvement is enabled by H100's Transformer- Engine FP8 path. Sibling to the prior 2026-05-05 Pantheon/Hydra, 2026-05-06 Spark Connect / Serverless Compute, and 2026-05-07 Lakebase FPW-elimination ingests, completing a four-article Databricks platform-engineering arc covering observability + OLAP-compute + OLTP + managed-inference altitudes. -
2026-05-07 — sources/2026-05-07-databricks-how-lakebase-architecture-delivers-5x-faster-postgres-writes (Tier-3 Databricks Engineering post — fifth canonical Lakebase ingest + first mechanism-level disclosure of the pageserver's internals beyond the name-level framing that prior Lakebase sources established; Databricks' third architectural-retrospective in three days, extending the 2026-05 Databricks engineering ingest window across both the OLTP (Lakebase) and OLAP (Serverless Compute) altitudes). Thesis: classical Postgres's Full Page Write (FPW) primitive — which copies an entire 8 KB page into WAL the first time it's modified after each checkpoint so recovery can tolerate torn pages — is structurally redundant when compute is stateless and streams WAL to a Paxos-based safekeeper quorum (first wiki-explicit disclosure of the safekeeper's durability primitive). Verbatim: "Because there is no local-disk page to tear, the failure mode FPW was designed to prevent simply does not exist." But FPW had an incidental read-path role: its periodic images bounded delta-chain replay on the pageserver. Databricks solved this by pushing image generation down to the storage tier — the pageserver generates images when a page has accumulated more delta records than a configured threshold without an intervening image; compute sends only compact deltas. Three benefits named: network efficiency (94% WAL reduction), scalability (image generation shared across multiple pageservers in the background), optimal reads (per-page-change-rate cadence, not checkpoint-scoped). Quantified on HammerDB TPROC-C: 4 vCPU +20%, 16 vCPU 2.8×, 32 vCPU 4.5×+ (95,686 → 439,300 NOPM); WAL/transaction 58 KB → <4 KB (94% reduction); the pre-change flat 16v→32v scaling (95,832 → 95,686) canonicalises "compute resources were not being used because FPW was the bottleneck". Production customer (56 vCPU): WAL rate 30 MB/s → 1 MB/s (30× reduction). Read-path dividend: p99 −30% to −50%, p50 ~−30%. Synced Tables ingestion 17k → 62k rows/sec (3.6×). Rollout: "since late March" → globally active 2026-05-07 (~6-week window) via the existing Postgres
XLOG_FPW_CHANGEWAL record mechanism — canonicalised as the live-wal-protocol-switch-via-xlog-fpw-change pattern (in-log feature flag using a pre-existing control record that both compute and storage already parse; zero customer restarts). New canonicalisations (8 new wiki pages): 4 concepts (postgres-full-page-write, torn-page, concepts/wal-write-ahead-logging, delta-chain-replay) + 2 patterns (image-generation-pushdown-to-storage, live-wal-protocol-switch-via-xlog-fpw-change) + 1 system (systems/hammerdb) + the source page. Extended (5 existing pages): systems/lakebase (new axis section + Seen-in entry), systems/pageserver-safekeeper (image-generation responsibility + Paxos-quorum framing + XLOG_FPW_CHANGE rollout mechanism), systems/postgresql (first wiki quantification of 15×-WAL-inflation FPW ceiling +XLOG_FPW_CHANGEas live-rollout vehicle), concepts/compute-storage-separation (fifth axis — enabler for structural elimination of durability primitives, not just relocation), concepts/wal-write-ahead-logging (WAL also carries control records that enable atomic protocol-switch rollouts). This ingest completes Lakebase's five-axis canonicalisation: CMK encryption (2026-04-20) + bursty workloads (2026-04-27) + agent provisioning (2026-04-29) + PITR operations (2026-04-30) + performance engineering via storage-work-offload (2026-05-07). Companion post: "Zero- downtime patching: Lakebase Part 1 — prewarming" (earlier 2026-05 Databricks post on cache prewarming); common thread is "move heavy-lifting tasks away from your transactions and into our scalable background storage stack." - 2026-05-06 — sources/2026-05-06-databricks-rethinking-distributed-systems-for-serverless-performance (Tier-3 Databricks Engineering post — the second architectural-retrospective in two days extending the 2026-05 Databricks engineering ingest window). Thesis quote: "Stability becomes a system property rather than a user responsibility, enabled by architectures that isolate workloads, intelligently place them, and dynamically adapt resources." Three systems compose into Databricks Serverless Compute: (1) Spark Connect — the gRPC client-server rearchitecture of Spark's driver model, framed as "the most significant architectural transformation in Spark's history". User application code no longer co-executes with the driver; queries travel as serialised logical plans. Unit of execution shifts from processes to queries. Canonical instance of client-server-decoupling + grpc-decoupled-driver-client. Enables Databricks' disclosed 25+ major Spark runtime upgrades per year with 99.998% success rate across >2 billion workloads (cited from SIGMOD/PODS '25 Breese et al. paper "Blink Twice: Automatic Workload Pinning and Regression Detection for Versionless Apache Spark using Retries"). (2) Serverless Gateway — workload-aware router combining three real-time signals (query size from logical plan + cluster utilisation + interactive-vs-batch latency profile) with continuous re-evaluation as conditions shift. Canonicalises patterns/ai-gateway-provider-abstraction and resolves the utilization-vs-predictability-tradeoff at the pool layer. (3) Serverless Autoscaler — two-axis adaptive autoscaler scaling "horizontally and vertically" (concepts/elasticity); OOM-aware VM-restart primitive (adaptive-oom-recovery + oom-aware-vm-restart-autoscaling) detects task OOM and restarts the task on a larger VM without job failure. Customer outcomes: **CKDelta 20 min vs 4–5 hr (12–15× speedup), Unilever 2–5× faster + 25% cost reduction, HP 32% cloud savings
- 36% runtime reduction. Paired with the 2026-05-05 Pantheon/ Hydra ingest, this forms the 2026-05 Databricks architecture double-ingest** breaking the late-April / early-May Tier-3 Databricks marketing skip streak (23 new wiki pages touched: 1 source + 4 systems + 6 concepts + 3 patterns + extensions to apache-spark + databricks + companies/databricks).
- 2026-05-05 — sources/2026-05-05-databricks-10-trillion-samples-a-day-scaling-beyond-traditional-monitoring (Tier-3 Databricks Engineering post — the canonical architectural-retrospective that broke the Databricks 2026-05 skip streak). Scale datum: ~70 cloud regions / 3 major clouds / 160+ Pantheon instances / 5B active in-memory timeseries / 10T samples/day / 20B unaggregated timeseries in Hydra. Three coupled architectural responses to an order-of-magnitude growth in monitoring load: (1) Pantheon — a Thanos fork scaled to hyperscale with two Receive groups on distinct memory- retention tiers (2h / 30m), three isolated StatefulSets per group preserving quorum with stronger operational isolation, at-least-once block uploads (2 of 3 StatefulSets), and a purpose-built 3-controller control plane (Rollout Operator / Hashring Controller / Autoscaling + Self-Healing Controller) running "dozens of automations per week". (2) A Telegraf + Dicer aggregation shield sustaining >1 GB/s per region across thousands of rules, absorbing a 2-5× incident surge so Pantheon only saw 20%. Canonical instance of sticky-routing-for-aggregator-state (Kafka rejected for cost + latency). (3) Hydra — a lakehouse-native platform for raw high-cardinality troubleshooting data: Spark Structured Streaming + Auto Loader → Delta Lake with ~5 min freshness and 50× cheaper storage than Thanos, queryable from Grafana via a PromQL-to-SQL translation layer (canonical promql-to-sql-over-delta-tables) — so engineers' dashboards keep working unchanged while the substrate shifts fundamentally. Three canonical concepts promoted: concepts/metric-cardinality (primary TSDB scaling factor), tsdb-scaling-bottleneck (scale-ups as daily events), serverless-workload-churn-cardinality (tens of millions of VMs daily drives label churn). Migration outcomes: "millions of dollars in annual cloud costs" saved, "~5× reduction in monitoring infrastructure downtime," many sources of manual toil eliminated. Unified metric semantics across Pantheon + Hydra — "engineers should not need to understand our ingestion architecture" — canonicalises the patterns/telemetry-to-lakehouse pattern at hyperscale. First canonical wiki instance of CUJs applied to observability-platform design.
- 2026-04-30 — sources/2026-04-30-databricks-backstage-with-lakebase
(Tier-3 Thoughtworks guest-post on Databricks Blog, Part
1 of a three-part series — Part 2: Governance, Part 3:
FinOps — forthcoming). Fourth canonical wiki source on
Lakebase after CMK (2026-04-20),
LangGuard (2026-04-27), and Stripe Projects (2026-04-29).
Thoughtworks runs a proof-of-concept ripping
Backstage (Spotify's state-heavy
Internal Developer Portal, endorsed on Thoughtworks'
Technology
Radar) off its standard Postgres database and pointing
it at Lakebase. Central thesis: when branching + PITR
become effectively free, two separate engineering
practices collapse into the same primitive
("branching is just PITR with source_branch_time =
now"), rearranging the developer cycle enough to
deprecate 20-30% of test code (mock objects). Load-
bearing operational disclosures: (a) Wire-protocol-
Postgres compatibility — "Because it speaks wire-
protocol Postgres, Backstage doesn't know or care that it
isn't talking to RDS"; Backstage's Knex migrations ran
cleanly, only
PgSearchEnginehad to be swapped for Backstage's default in-memory search. (b) Auth was the friction point — Lakebase rejects classic Databricks Personal Access Tokens and expects an OAuth JWT minted bydatabricks postgres generate-database-credential. Thoughtworks wrapped the command in a 50-minute cron rewritingDATABRICKS_TOKENin.env— canonical credential-refresh-cron-as-auth-compat-shim. (c) First wiki disclosure of Lakebase branching throughput at MB-scale dataset granularity: a 63 MB Backstage catalog branch lands in 1.09 seconds data plane (control-plane ack was instant). Prior Lakebase sources disclosed only "seconds" (LangGuard) or "sub-350 ms" cold Postgres provisioning (Stripe Projects); this is the first wiki measurement separating control-plane ack from data-plane clone. (d) First wiki disclosure of Point-in-Time Recovery at Lakebase altitude: wipe offinal_entities(32 rows → 0), then recovery branch from a pre-wipe timestamp, end-to-end in 3.78 seconds. Production still at zero during recovery (branches fully isolated). Canonical concepts/point-in-time-recovery. (e) WAL-record granularity disclosed: requested 22:56:02Z, got 22:55:50Z (12 seconds earlier) — PITR snaps backward to the nearest WAL record. Canonical concepts/wal-write-ahead-logging as an "important caveat for time-sensitive recovery workflows." (f) Architectural unification: branching-is-pitr-with-time-now — same control-plane call, same storage substrate, same compute-attach step; only the time parameter differs. Latency envelopes confirm (1.09 s vs 3.78 s). (g) Branch API gotcha: request body must nest everything inside aspecobject + explicitly specifyttl,expire_time, orno_expiry— branches are short-lived by default. (h) Developer-cycle transformation thesis: paired with the numbers, the POC argues cheap branching deprecates 20-30% of test code (mock objects — "not test coverage, that's test infrastructure") and shifts schema- migration validation from staging-deploy to development. See mock-object-maintenance-cost + concepts/integration-tests-against-real-database + patterns/database-branch-per-test-over-mocking. The load-bearing claim is the cost-benefit flip: "When branching a production-equivalent database costs nothing, mocking becomes the expensive choice." Introduces two new systems (systems/backstage, systems/databricks-postgres-cli), one MVP system (systems/thoughtworks-technology-radar), three new concepts (concepts/point-in-time-recovery, concepts/wal-write-ahead-logging, mock-object-maintenance-cost), two additional concepts (concepts/short-lived-credential-auth, concepts/integration-tests-against-real-database), and three new patterns (branching-is-pitr-with-time-now, patterns/database-branch-per-test-over-mocking, credential-refresh-cron-as-auth-compat-shim). Extends systems/lakebase, concepts/database-branching, concepts/copy-on-write-storage-fork, concepts/compute-storage-separation. Caveats: Tier-3 single-vendor POC with Thoughtworks as guest author; 1.09-s and 3.78-s numbers are single-shot in a development environment, not production-scale benchmarks; 63 MB is a developer-IDP-scale dataset (scaling behaviour at GB/TB not disclosed); 12-second WAL snap- back is a function of the POC's write cadence, not a general guarantee; 50-minute cron refresh is a POC hack, not a production pattern; 20-30% mock-code claim is attributed to "multiple partner teams" without methodology; Part 1 of 3 — governance + FinOps content forthcoming. Cross-source continuity: fourth Lakebase source on the wiki; together the four sources now canonicalise Lakebase's compute-storage-separated architecture at four distinct altitudes — encryption (CMK), bursty-workload fit (LangGuard), agent-lifecycle substrate (Stripe Projects), developer-cycle transformation (Backstage). Tier-3 on-scope rationale: architecture + mechanism + numbers density >60% — explicit dataset-size + wall-clock numbers (63 MB, 1.09 s, 3.78 s, 12 s snap-back, 50-min cron), explicit mechanism disclosure (copy-on-write pointer semantics, WAL-record snap-back, branching-≡-PITR unification, OAuth-JWT vs PAT auth posture,spec-nested API with explicit lifetime), explicit integration trade-off discussion (mock deprecation, schema-migration-validation shift). Passes the 20% borderline-case threshold decisively. Not a product-launch post; it's a workflow-transformation retrospective by a consultancy.) - 2026-04-29 — sources/2026-04-29-databricks-and-stripe-projects-infrastructure-built-for-agents
(Tier-3 short joint launch post co-bylined by Brad Van
Vugt and Guillaume Rivals announcing Databricks as a
launch partner for Stripe
Projects — the agent-first CLI Stripe launched
2026-04-30. The Databricks side of the integration
exposes Neon Postgres databases (under the
Lakebase architecture name) as
agent-provisionable resources through the Stripe Projects
catalog, making Databricks the second launch-side
provider in the
agent-provisioning protocol after Cloudflare. Disclosed
operational datum: <350 ms for an agent-driven
production-ready Neon Postgres via the CLI — the first
wiki operational number for the protocol's per-request
latency envelope at the database-resource tier. Collapses
the prior Lakebase-vs-Neon distinction by naming both
interchangeably ("Lakebase architecture, developed by
Neon" + "bringing Neon databases seamlessly"). The
post articulates the three architectural pillars that
make agent-driven OLTP viable on this substrate:
serverless scale-to-zero, instant
database branching via
zero-copy cloning
for "safely test code, run migrations, or experiment
with new prompts against live data states", and
Postgres compatibility as an agent-ergonomic property
("agents understand Postgres better than any other
OLTP database"). Third canonical cross-source
confirmation of concepts/compute-storage-separation
as Lakebase's load-bearing property — new axis:
per-request compute lifecycle at agent-initiated cadence
("agents can create, build, and tear down OLTP
databases in seconds"). Introduces
agent-provisioned-database as a new concept
(database-tier sibling of
agent-provisioned-account) and adds Lakebase
as third known-use of
patterns/partner-managed-service-as-native-binding
(after Cloudflare/PlanetScale + Fly.io/Tigris) — first
agent-as-customer instance of that pattern, with a
payments-platform-orchestrator rather than a compute-
platform-orchestrator. Co-announces but does not deep-
dive the separate Stripe Data Pipeline × Databricks
Marketplace zero-ETL integration (deferred to sibling
post "Stripe data now available in Databricks"). Short
post, high marketing density; Tier-3 ingest-threshold
passed on (a) architectural density of three-pillar
paragraph >20%, (b) new operational datum <350 ms, (c)
first wiki record of second launch-partner in the
protocol, (d) new
agent-provisioned-databaseconcept. Caveats: spend-cap / rate-limit / fraud-heuristic / orphan-cleanup policies not disclosed; <350 ms is typical-case single-shot, not concurrent-burst; Lakebase ↔ Neon naming collapsed; no internal-mechanism disclosure on any of the three architectural pillars beyond verbatim quotes. Introduces agent-provisioned-database. Extends systems/lakebase, systems/stripe-projects, concepts/scale-to-zero, concepts/compute-storage-separation, concepts/database-branching, concepts/copy-on-write-storage-fork, agent-provisioning-protocol, patterns/partner-managed-service-as-native-binding. Cross-source continuity: companion-launch-partner to sources/2026-04-30-cloudflare-agents-can-now-create-cloudflare-accounts-buy-domains-and-deploy (one-day earlier Cloudflare + Stripe launch of the protocol); third Lakebase source on the wiki after sources/2026-04-20-databricks-take-control-customer-managed-keys-for-lakebase-postgres|2026-04-20 CMK and sources/2026-04-27-databricks-inside-one-of-the-first-production-deployments-of-lakebase-langguard|2026-04-27 LangGuard; collectively these three sources canonicalise Lakebase's compute-storage-separated architecture at three distinct altitudes — encryption (CMK), bursty- workload fit (LangGuard), agent-lifecycle substrate (this post).) - 2026-04-29 — sources/2026-04-29-databricks-approximate-answers-exact-decisions-new-sketch-functions-for-analytics (Tier-3 Databricks product-engineering post announcing four new Apache DataSketches-backed sketch function families in Databricks SQL / DataFrame / Structured Streaming: KLL for percentiles, Theta for distinct-count with set algebra, Approximate top-K for heavy hitters, Tuple for distinct
- metric aggregation. Canonicalises decision-support vs audit query as the architectural classifier that determines whether it's safe to accept 1–2% relative error in exchange for orders-of-magnitude compute reduction. Framing: "If knowing '~4.7M unique users ±1%' leads to the same decision as '4,712,389 unique users,' the approximate answer at a fraction of the cost is strictly better." The pattern contribution is sketch as Delta BLOB column — build once at ETL, merge on read in milliseconds — extending the pre-existing sketch-as-BLOB pattern (PlanetScale Insights) to the Delta Lake substrate and exposing it at the SQL aggregate level. Named concrete use cases: latency monitoring dashboards ("168 precomputed sketches" for a weekly trending view), audience-overlap marketing measurement via Theta union/intersection/ difference (Super Bowl ad ∩ Instagram campaign), live clickstream leaderboards via approximate-top-K streaming micro-batch merges, unique-customer-plus-revenue composed aggregations via Tuple sketches. The post is explicit about the boundary: "When to stay exact: Financial auditing, compliance reporting, or any use case where regulatory or business requirements demand precise values." Critically positions mergeability — the associative merge operator — as the architectural unlock, turning sketches from "a faster percentile" into a storage primitive that converts dashboards from scans into reduces. Community contribution: Christopher Boumalhab (cboumalh on GitHub) implemented the Theta and Tuple sketch function families in upstream Apache Spark. Agent-surface continuity: Genie Code can recommend the right sketch family — Databricks' ongoing pattern of threading agentic surfaces through every new capability. Introduces systems/apache-datasketches, concepts/probabilistic-data-structure, concepts/probabilistic-data-structure, concepts/probabilistic-data-structure, concepts/probabilistic-data-structure, concepts/probabilistic-data-structure, decision-support-vs-audit-query, precomputed-sketch-column-in-delta-table, set-algebra-on-theta-sketches. Extends concepts/probabilistic-data-structure, ddsketch-error-bounded-percentile, sketch-as-mysql-binary-column, local-global-aggregation-decomposition. Fifth Databricks-platform face on the wiki (after Redpanda/Iceberg analytics, Santander integration, Zalando Spark workload, multimodal-substrate, DICER auto-sharder/MLflow/Lakebase, and AutoCDC declarative CDC) and canonicalises the Databricks-as- approximate-analytics-primitive-vendor framing. Vendor post; benchmark claims self-reported ("1000× speedup", "orders of magnitude less compute") and not independently verified.)
- 2026-04-27 — sources/2026-04-27-databricks-inside-one-of-the-first-production-deployments-of-lakebase-langguard (Tier-3 Databricks case study profiling LangGuard — one of the first startups building its production governance engine on systems/lakebase. Introduces systems/langguard as a runtime enforcement layer for enterprise agentic workflows and GRAIL as its patent-pending live-knowledge-graph governance fabric. Canonicalises the three-property Lakebase fit for bursty agentic workloads: (1) serverless autoscaling + concepts/scale-to-zero between bursts — agent workflows are dormant for hours, then hundreds of trace writes + enforcement reads in seconds; (2) millisecond reads via compute-local cache for hot indexed lookups against GRAIL context + policy tables keeps governance off the agent's critical path; (3) instant copy-on-write database branching for testing new governance policies against real production trace data in seconds without risking the live environment — the first canonical wiki instance of DB branching at governance-policy-validation altitude (distinct from schema-change-testing and dev-sandbox uses). LangGuard's team previously built IBM QRadar (SIEM at petabytes/day); the QRadar-era lesson cited verbatim: "database architecture is destiny" for bursty security-telemetry workloads; coupled-compute/storage Postgres forced provisioning for peak and paying for idle around the clock. Stated roadmap: train behavioral baselines on historical GRAIL trace data via MLflow to move from reactive runtime enforcement to predictive governance; co-location of operational trace data with the analytical platform removes the ETL barrier that normally separates the two. Enterprise workflow scope cited: "tens of coordinated agents, hundreds of tool invocations, multiple foundation models, and policies managed across fifteen or more enterprise Systems of Record" — ServiceNow, IAM/IDP, Salesforce, Workday, Wiz, CrowdStrike, TalkDesk, MCP Gateways, API Gateways. Introduces systems/langguard, systems/grail-data-fabric, concepts/governed-agent-data-access, concepts/attribute-based-access-control, agent-behavioral-baseline, runtime-governance-enforcement-layer, patterns/database-branch-per-test-over-mocking; extends systems/lakebase, concepts/compute-storage-separation, concepts/database-branching, concepts/copy-on-write-storage-fork, concepts/scale-to-zero, bursty-query-pattern. Tier-3 Databricks — borderline product-case-study post (joint vendor narrative with LangGuard), ingested because architectural content is substantive (~60% of body: bursty-workload shape, compute/storage-separation payoff for scale-up with no data movement, copy-on-write branching for policy testing, QRadar lineage as empirical prior) even though ~40% is product positioning. No concrete latency/throughput numbers disclosed — GRAIL schema + query patterns undisclosed; predictive governance is roadmap, not shipped.)
- 2026-04-22 — sources/2026-04-22-databricks-stop-hand-coding-change-data-capture-pipelines
(Tier-3 Databricks product-engineering post pitching
AutoCDC inside
Lakeflow SDP
as the declarative replacement for hand-rolled
MERGElogic in CDC / SCD pipelines. Names the four structural pain sources of hand-rolled CDC (out-of-order updates, duplicate events, delete application, idempotency across retries) and maps each to a declarative API parameter (keys,sequence_by,apply_as_deletes,stored_as_scd_type). Canonicalisessequence_byas the load-bearing primitive for out-of-sequence CDC event handling — a concern separate from idempotency that hand-rolled pipelines routinely get wrong (arrival-order-as-logical-order assumption, dedup without ordering, window-function dedup without retry-safety). Three input modes supported: native Change Data Feed (CDF), CDF with SCD Type 2 history, and snapshot-diff inference (the third canonical CDC ingest shape on this wiki after log-based capture and CDF). Code footprint disclosure: 6–10 lines AutoCDC vs 40–200+ lines hand-rolled MERGE; Fortune 500 Aerospace & Defense adopter quote: "4 lines of code could replace what I was doing in 1,500 lines of code before." Databricks Runtime performance improvements since Nov 2025: 71% better perf-per-dollar on SCD Type 1, 96% on SCD Type 2 — propagated universally to all AutoCDC pipelines because the declarative API lets Runtime-level optimisations apply across the fleet. Named regulated-vertical adopters: Navy Federal Credit Union (billions of events/day real-time event processing), Block (pipeline dev time: days → hours), Valora Group (Swiss foodvenience retail master-data CDC). Rhetorical framing: Databricks argues LLM codegen does not solve CDC correctness ("LLMs can generate code, but they don't understand your data") and positions Genie Code as the AI codegen client that produces AutoCDC declarations rather than rawMERGE— LLM correctness envelope becomes equal to the declarative API's bounded envelope. Reinforces MERGE over INSERT OVERWRITE as the runtime primitive — AutoCDC displaces hand-authoring, not MERGE itself. Introduces systems/databricks-autocdc, systems/databricks-genie-code, concepts/change-data-capture, concepts/change-data-capture, declarative-cdc-over-hand-rolled-merge. First wiki source on declarative CDC as a distinct pattern axis separate from the CDC-ingest-mode taxonomy (log-based vs CDF vs snapshot-diff). Tier-3 Databricks — ingested because architectural framing is substantive (≥50% of body is declarative-vs-hand-rolled tradeoff + API-parameter semantics - before/after code + perf / code-reduction numbers + named adopters); fails no skip signals.)
- 2026-04-22 — sources/2026-04-22-databricks-multimodal-data-integration-production-architectures-for-healthcare-ai
(Tier-3 Databricks healthcare-vertical post that is
architecturally non-trivial: names the specialty-store-per-
modality failure mode (FHIR store + omics store + imaging
store + vector store — duplicated governance, brittle
cross-store joins) and positions the lakehouse-as-multimodal-
substrate as its remedy. Every modality (genomics, imaging
features, clinical-notes entities, wearables aggregates)
lands in governed Delta tables under
one Unity Catalog governance
surface; modality-specific tooling —
Glow (VCF / BGEN / PLINK → Delta),
Mosaic AI Vector Search
over imaging-derived feature embeddings,
Lakeflow SDP
(
@dp.table+@dp.materialized_viewfor wearables streaming) — sits above the substrate rather than beside it. Canonicalises the four-fusion-strategy taxonomy (early-fusion / intermediate-fusion / late-fusion / attention-based-fusion) paired with deployment-reality triggers — "match fusion to your deployment reality: modality availability patterns, dimensionality balance, and temporal dynamics" — and names the missing-modality problem ("missingness isn't an edge case — it's the default") with three production responses (modality masking during training, sparse / modality-aware attention, transfer learning). Reproducibility story pinned to Delta time travel + CI/CD + MLflow experiment tracking. Introduces systems/databricks-glow, systems/lakeflow-spark-declarative-pipelines, systems/mosaic-ai-vector-search, governed-delta-tables-per-modality, fusion-strategy-selection-by-deployment-reality, early-fusion, intermediate-fusion, late-fusion, attention-based-fusion, missing-modality-problem, modality-masking-during-training. No production metrics disclosed (vendor post), no code snippets. Extraction scoped to the architectural content; healthcare-vertical specifics — tumor boards, trial matching, 28 CFR Part 202 tagging — preserved only where they illustrate a sysdesign point. First wiki ingest naming Lakeflow SDP, Glow, Mosaic AI Vector Search.) - 2026-04-22 — sources/2026-04-22-databricks-are-llm-agents-good-at-join-order-optimization
(Databricks + UPenn research prototype applying a frontier
LLM agent as an offline
join-order tuner for the Databricks query engine. Names
the three-component optimizer decomposition (cardinality
estimator + cost model + search) and frames the LLM as
offline-only — too slow for the optimizer hot path, but
the perfect fit for the historically-human DBA tuning loop
(offline-query-tuning-loop). Architecture:
single tool
execute_plan(candidate)returning runtime + subplan sizes; rollout budget (50 prototype, 15 eval); grammar-constrained structured output admitting only valid join reorderings (structured-output-grammar-for-valid-plans); best-of-N selection (anytime-optimization-algorithm / rollout-budget-anytime-plan-search). Evaluation on JOB (113 queries, 10× IMDb): 1.288× geomean / 41% P90 — beating perfect cardinality estimates (because measurement beats estimation), smaller LLMs, and classical BayesQO Bayesian-optimization baseline (Postgres-tuned asymmetry noted). Canonical win: JOB query 5b'sLIKE-predicate failure case (like-predicate-cardinality-estimation-failure). Introduces systems/databricks-join-order-agent, systems/join-order-benchmark-job, systems/bayesqo, join-order-optimization, cardinality-estimation, llm-agent-as-query-optimizer, anytime-optimization-algorithm, offline-query-tuning-loop, like-predicate-cardinality-estimation-failure, exploration-exploitation-tradeoff-in-agent-search, llm-agent-offline-query-plan-tuner, structured-output-grammar-for-valid-plans, rollout-budget-anytime-plan-search. Tier-3 Databricks — borderline research/ML post, but content is substantively about database-engine optimizer architecture (join ordering, cardinality estimation, cost models, plan search) rather than ML methodology. Ingested for the architectural framing. Research prototype, not a shipping feature; cost-per-tuned-query not disclosed.) - 2026-04-17 — sources/2026-04-17-databricks-governing-coding-agent-sprawl-with-unity-ai-gateway (Launch post for Unity AI Gateway coding-agent support. Names coding- agent sprawl — engineers routinely mix Cursor + Codex + Claude Code + Gemini CLI + others in parallel — as the forcing function. Three-pillar answer: (1) centralised security + audit via Unity Catalog + MLflow tracing + single-SSO across all coding tools + Databricks-managed MCP servers; (2) single bill + cost controls via Foundation Model API first-party inference + BYO external capacity + per-developer (not per-tool) budgets — patterns/unified-billing-across-providers; (3) full observability via OpenTelemetry → Unity-Catalog-managed Delta tables joinable with HR/PR-velocity data — patterns/telemetry-to-lakehouse. Launch-day clients: Cursor, Codex CLI, Gemini CLI; Claude Code via MLflow 3 tracing. Structural mirror of the Cloudflare internal AI engineering stack shape with different substrates. Extends the wiki's existing AI-gateway provider- abstraction pattern along two new axes: coding-tool clients as first-class, and MCP-server governance as a peer concern to LLM-call governance. Introduces systems/unity-ai-gateway, systems/databricks-foundation-model-api, systems/cursor, systems/claude-code, systems/codex-cli, systems/gemini-cli, coding-agent-sprawl, concepts/centralized-ai-governance, patterns/central-proxy-choke-point, patterns/telemetry-to-lakehouse, patterns/unified-billing-across-providers. Tier-3 Databricks — ingested for framing + integration architecture, not internals; product-announcement post, gateway routing/fallback/rate-limiter mechanics not disclosed.)
- 2026-04-20 — sources/2026-04-20-databricks-take-control-customer-managed-keys-for-lakebase-postgres (Customer-Managed Keys rollout for systems/lakebase — Databricks' serverless Neon-descended Postgres. Three-level concepts/envelope-encryption hierarchy CMK → KEK → DEK; CMK held in customer's cloud KMS (systems/aws-kms / systems/azure-key-vault / systems/google-cloud-kms); cryptographic-shredding on revocation across both persistent (Pageserver+Safekeeper) and ephemeral (Postgres compute VM) layers; per-boot-ephemeral-key pattern for VM-local state; seamless key rotation as a property of the envelope hierarchy; Account↔Workspace delegation for separation of duties; auditability in customer KMS tenancy. Enterprise tier. Introduces systems/lakebase, systems/pageserver-safekeeper, systems/aws-kms, systems/azure-key-vault, systems/google-cloud-kms, concepts/envelope-encryption, concepts/byok-bring-your-own-key, cryptographic-shredding, per-boot-ephemeral-key.)
- 2026-04-20 — sources/2026-04-20-databricks-mercedes-benz-cross-cloud-data-mesh
(Mercedes-Benz cross-cloud data mesh on Unity Catalog + Delta
Sharing + Delta Deep Clone; AWS Iceberg-on-Glue producer ↔ Azure
Delta-on-ADLS consumers; hybrid "live share vs. incremental replica"
tier; 66 % egress cost reduction on first 10 data products, ~93 %
projected annual at 50 use cases; weekly-load → every-second-day
freshness; DDX self-service orchestrator, DABs + Azure DevOps
deploys, Sync-Job-bytes → producer chargeback,
VACUUMfor GDPR delete propagation on replicas; introduces systems/delta-sharing, systems/delta-lake, systems/mercedes-benz-data-mesh, concepts/data-mesh, concepts/egress-cost, hub-and-spoke-governance, cross-cloud-architecture, cross-cloud-replica-cache, chargeback-cost-attribution.) - 2026-01-13 — sources/2026-01-13-databricks-open-sourcing-dicer-auto-sharder (open-sourcing Dicer, Databricks' auto-sharder; dynamic slice-range sharding with hot-key isolation/replication, eventually-consistent Assignments, state transfer across reshards; positioned vs. prior art Slicer / Centrifuge / Shard Manager; three production case studies — Unity Catalog 90–95% hit rate, SQL Query Orchestration zero-downtime scaling, Softstore 85% hit rate across rolling restarts via state transfer; use cases include LLM KV cache / LoRA-adapter GPU placement, batch aggregation, soft leader selection, rendezvous coordination.)
- 2025-12-03 — sources/2025-12-03-databricks-ai-agent-debug-databases (internal AI agent platform Storex for DB debugging across thousands of instances / 3 clouds / hundreds of regions / 8 regulatory domains; central-first sharded foundation; DsPy-inspired tools-as-functions framework; snapshot-replay validation with judge LLM; specialized per-domain agents; hackathon → platform journey; claimed up to 90% investigation-time reduction, <5 min new-hire ramp-up.)
- 2025-10-01 — sources/2025-10-01-databricks-intelligent-kubernetes-load-balancing (proxyless client-side L7 LB + custom xDS EDS; P2C + zone-affinity with spillover; rejected Istio / headless services; 20% pod-count reduction; surfaced cold-start problem that long-lived L4 LB had hidden.)
Ingest posture¶
Tier-3 filter applies: by default skip product PR, acquisition news,
pure ML methodology posts. Ingest when the article covers:
distributed-systems internals, scaling trade-offs, Kubernetes / network
infrastructure, production incidents, storage/streaming design, or
data-platform internals (Photon, Delta Lake, Unity Catalog — when
architecturally substantive). Several 2025 posts already reviewed and
logged as off-topic in log.md (TAO LLM-tuning, Neon acquisition PR,
Data Intelligence for Marketing launch).