Skip to content

SYSTEM Cited by 1 source

Basin Catalog

Basin Catalog is Cloudflare's managed Apache Iceberg REST catalog — the storage / metadata layer of Basin. It was the first product Cloudflare launched in the family (at Birthday Week 2025, then named R2 Data Catalog); it reached GA and was renamed on 2026-10-01. Documented at developers.cloudflare.com/basin-catalog. (Source: sources/2026-10-01-cloudflare-introducing-cloudflare-basin-an-open-serverless-data-platform)

See also R2 Data Catalog — the prior name, canonicalised earlier via the internal Town Lake post, where it serves as that platform's cold/warm Iceberg tier. Basin Catalog is the same product, rebranded.

What it is

"The easiest way to get started with Apache Iceberg." One command — npx wrangler basin catalog create CATALOG_NAME — yields "a fully managed Apache Iceberg REST catalog that automatically performs routine maintenance required to keep those tables performant and healthy." Tables live as Parquet files on R2 under Iceberg's open table format; Cloudflare operates the catalog so developers don't run their own Iceberg metastore.

Use spans "simple use cases such as giving DuckDB a structured way to access analytics data in R2, all the way to developers building complete enterprise data sharing platforms, fully taking advantage of zero egress fees and easy-to-use APIs." Thousands of developers since launch.

Automatic table maintenance

The core value proposition: Basin Catalog absorbs the OTF externalisation cost — the compaction / GC / metadata loop that Iceberg otherwise pushes into customer code. At launch it had just added automatic compaction; since then:

  • Per-table compaction policies — choose target file sizes based on each table's access pattern (compaction of snapshot-fragmented small files into larger ones).
  • Automatic snapshot expiration — remove old Iceberg snapshots per a retention policy while preserving a minimum number of recent snapshots.
  • Unreferenced data-file cleanup — reclaim storage when snapshots expire, "without requiring a separate Spark maintenance job."
  • Manifest optimization — consolidate + cluster fragmented manifests by partition before compaction, reducing metadata I/O during query planning.

The planning-statistics hand-off to Basin SQL

Compaction isn't just housekeeping — it feeds the query engine. "As datasets grow, Basin Catalog compacts metadata and data files to reduce I/O and generates statistics for query planning. Basin SQL uses those statistics to split queries into smaller tasks and distribute them across Workers." The catalog's maintenance output is the input to Basin SQL's distributed scatter-gather planning.

Roadmap (stated)

  • A new way for compaction to "efficiently sort and cluster data for improved query performance."
  • More granular auth controls for namespaces and tables.
  • Jurisdiction support to adhere to data sovereignty and compliance requirements.

Caveats

  • GA announcement — no compaction throughput / latency / cost numbers.
  • The relationship between Basin Catalog's maintenance cadence and Iceberg-metadata-size pressure at scale is not quantified.

Seen in

Last updated · 766 distilled / 2,225 read