SYSTEM Cited by 1 source
Basin Catalog¶
Basin Catalog is Cloudflare's managed Apache Iceberg REST catalog — the storage / metadata layer of Basin. It was the first product Cloudflare launched in the family (at Birthday Week 2025, then named R2 Data Catalog); it reached GA and was renamed on 2026-10-01. Documented at developers.cloudflare.com/basin-catalog. (Source: sources/2026-10-01-cloudflare-introducing-cloudflare-basin-an-open-serverless-data-platform)
See also R2 Data Catalog — the prior name, canonicalised earlier via the internal Town Lake post, where it serves as that platform's cold/warm Iceberg tier. Basin Catalog is the same product, rebranded.
What it is¶
"The easiest way to get started with Apache Iceberg." One command —
npx wrangler basin catalog create CATALOG_NAME — yields "a fully managed
Apache Iceberg REST catalog that automatically performs routine maintenance
required to keep those tables performant and healthy." Tables live as
Parquet files on R2
under Iceberg's open table format; Cloudflare operates the catalog so
developers don't run their own Iceberg metastore.
Use spans "simple use cases such as giving DuckDB a structured way to access analytics data in R2, all the way to developers building complete enterprise data sharing platforms, fully taking advantage of zero egress fees and easy-to-use APIs." Thousands of developers since launch.
Automatic table maintenance¶
The core value proposition: Basin Catalog absorbs the OTF externalisation cost — the compaction / GC / metadata loop that Iceberg otherwise pushes into customer code. At launch it had just added automatic compaction; since then:
- Per-table compaction policies — choose target file sizes based on each table's access pattern (compaction of snapshot-fragmented small files into larger ones).
- Automatic snapshot expiration — remove old Iceberg snapshots per a retention policy while preserving a minimum number of recent snapshots.
- Unreferenced data-file cleanup — reclaim storage when snapshots expire, "without requiring a separate Spark maintenance job."
- Manifest optimization — consolidate + cluster fragmented manifests by partition before compaction, reducing metadata I/O during query planning.
The planning-statistics hand-off to Basin SQL¶
Compaction isn't just housekeeping — it feeds the query engine. "As datasets grow, Basin Catalog compacts metadata and data files to reduce I/O and generates statistics for query planning. Basin SQL uses those statistics to split queries into smaller tasks and distribute them across Workers." The catalog's maintenance output is the input to Basin SQL's distributed scatter-gather planning.
Roadmap (stated)¶
- A new way for compaction to "efficiently sort and cluster data for improved query performance."
- More granular auth controls for namespaces and tables.
- Jurisdiction support to adhere to data sovereignty and compliance requirements.
Caveats¶
- GA announcement — no compaction throughput / latency / cost numbers.
- The relationship between Basin Catalog's maintenance cadence and Iceberg-metadata-size pressure at scale is not quantified.
Seen in¶
- sources/2026-10-01-cloudflare-introducing-cloudflare-basin-an-open-serverless-data-platform — canonical wiki source. GA + rename from R2 Data Catalog; four maintenance features; planning-statistics hand-off to Basin SQL.
Related¶
- systems/basin — the umbrella platform.
- systems/basin-pipelines — writes Iceberg tables into Basin Catalog.
- systems/basin-sql — consumes Basin Catalog's planning statistics.
- systems/apache-iceberg — the open table format managed here.
- systems/cloudflare-r2 — the underlying object-storage substrate.
- systems/cloudflare-r2-data-catalog — the prior name of this product.
- systems/apache-parquet — the columnar file format.
- concepts/open-table-format — the externalisation cost absorbed here.
- concepts/compute-storage-separation — the structural property.
- concepts/lsm-compaction — the compaction family.
- companies/cloudflare — operator.