SYSTEM Cited by 3 sources
DuckDB¶
DuckDB is an open-source, in-process columnar analytical database (duckdb.org) — designed for single-node analytical workloads, embedded into application processes the same way SQLite is, with full SQL support and columnar execution.
The defining properties for this wiki:
- In-process / embedded — no separate server; runs in the application's address space.
- Columnar storage + vectorised execution — analytical query performance comparable to dedicated OLAP engines on single-machine workloads.
- Open-format reader/writer — reads/writes Parquet, Delta, and Iceberg via extensions.
- Single-node first — no built-in distributed execution; scales by machine size, not cluster size.
Created in 2018 by Hannes Mühleisen and Mark Raasveldt, researchers from the CWI lab behind MonetDB/X100 — carrying that lineage's focus on extracting maximum performance from a single CPU via vectorized execution. Explicitly modeled on SQLite's embeddable-library form factor, generalized to analytics. The extreme placement compiles DuckDB to WebAssembly to run entirely in a browser tab (shell.duckdb.org). Warfield's framing: "the glibc of structured data."
Seen in¶
-
sources/2026-08-26-allthingsdistributed-duckdb-and-the-changing-physics-of-analytics — Canonical wiki disclosure of DuckDB's origin, thesis, and AWS acquisition. Andy Warfield (AWS) presents DuckDB as the flagship of the single-node analytics resurgence driven by changing hardware physics (~50× more memory/cores/network on a single instance vs. 2007). Positioned as the embedded analytics engine (analytics-engine-as-embedded-library) — SQLite for OLAP. AWS sponsored DuckDB's Iceberg extension (Iceberg v2 + v3, >800K downloads/week) to broaden Iceberg beyond the Spark landscape and make S3 Tables a great fit for DuckDB users; that work motivated async I/O to saturate the NIC while scanning S3 (DuckDB 2.0). DuckLabs joins AWS as a subsidiary; DuckDB stays open source (MIT) under the DuckDB Foundation, team remains in Amsterdam. Anchored intellectually in the COST paper.
-
sources/2026-05-14-databricks-expanded-interoperability-with-unity-catalog-open-apis — First wiki disclosure as a UC-integrated external engine. Named alongside Apache Spark and Apache Flink as one of three external engines that "can create and write to UC managed Delta tables with centralized governance and automatic optimizations" in the Beta. Composes with Delta Kernel (the Java/Rust library DuckDB leverages for UC-coordinated commits) and UC Credential Vending (the auth substrate). Specific role inside the article: DuckDB is the small-footprint single-node consumer in a list otherwise dominated by distributed cluster engines (Spark / Flink) — it represents the architectural class of "laptop or single-VM analyst doing managed-table writes against a governed catalog". Article framing: "thousands of customers use Unity Catalog to govern and access Delta Lake and Apache Iceberg tables, with dozens of integrations in the growing Unity Catalog ecosystem — from Apache Spark and Trino to DuckDB and Confluent Tableflow."
-
sources/2026-10-01-cloudflare-introducing-cloudflare-basin-an-open-serverless-data-platform — DuckDB as a named Iceberg-compatible engine for Basin. Cloudflare lists DuckDB (alongside PyIceberg, Snowflake, Apache Spark) as an engine that can "read and write your data" on Basin's Basin Catalog — "that kind of data portability is only possible with free egress." An early use is specifically "giving DuckDB a structured way to access analytics data in R2." Canonical datapoint for DuckDB as the single-node consumer in the storage-separated + open-table-format story — this time against a zero-egress object store (R2) rather than S3.
Related¶
- systems/delta-kernel — the protocol-abstraction library DuckDB uses for UC-coordinated Delta writes.
- systems/delta-lake — the table format DuckDB reads + writes.
- systems/unity-catalog — the catalog DuckDB integrates with via Open APIs.
- systems/uc-managed-tables — the managed-table primitive DuckDB can now write into (Beta, 2026-05-14).
- systems/apache-spark, systems/apache-flink — the two distributed external-engine peers in the same UC integration.
- concepts/open-table-format — architectural shape DuckDB participates in.
- systems/monetdb — the CWI columnar/vectorized ancestor DuckDB's creators came from.
- systems/sqlite — the embedded-library precedent DuckDB is modeled on.
- systems/apache-iceberg — the open table format DuckDB's AWS-sponsored extension reads/writes.
- systems/s3-tables — Iceberg-backed S3 tabular storage DuckDB targets.
- embedded-analytics-engine — the system class DuckDB defines.
- single-node-analytics-resurgence — the trend DuckDB leads.
- analytics-engine-as-embedded-library — DuckDB's core pattern.
- saturate-nic-with-async-io — DuckDB's S3/Iceberg scan technique.