Introducing Cloudflare Basin: an open, serverless data platform, now generally available¶
Summary¶
Cloudflare announced that the Cloudflare Data Platform (launched in open beta at Birthday Week 2025) is now generally available and rebranded as Basin — a serverless analytics data platform built on Apache Iceberg (the open table format standard) and R2 Object Storage. Basin is an end-to-end analytics platform covering ingestion, storage, and querying, organized around three now-GA products: Basin Pipelines (formerly Cloudflare Pipelines — receives events from Workers/HTTP/Logpush, transforms them with SQL, writes Iceberg tables or files to R2), Basin Catalog (formerly R2 Data Catalog — managed Iceberg REST catalog that auto-maintains tables for performance and cost), and Basin SQL (formerly R2 SQL — a serverless, distributed SQL engine that queries Iceberg tables directly on Cloudflare, splitting queries across Workers). The year-of-beta investment concentrated on three axes unique to analytics on the Developer Platform: speed, openness, and cost efficiency — the last anchored by R2's zero egress fees, which make the Iceberg open-table-format promise of "use the right query engine for the job" (PyIceberg, DuckDB, Snowflake, Apache Spark) economically practical.
Key takeaways¶
-
A basin is where rivers from many sources converge to a single point — the naming captures the architecture: Pipelines brings data in, Basin Catalog stores it, Basin SQL makes it instantly queryable. Today Basin spans ingestion + storage + querying and "will expand over time with more products managing the rest of the analytical data lifecycle." (Source: sources/2026-10-01-cloudflare-introducing-cloudflare-basin-an-open-serverless-data-platform)
-
The platform exists because of two 2024-era shifts. First, Apache Iceberg emerged as the standard open table format, making data "portable across nearly every major query engine." Second, developers started bringing analytics data to R2, where the lack of egress charges made it "practical and cost-efficient to actually access their data from different tools, teams, regions, and cloud providers." Basin is the platform built on top of those two shifts (concepts/open-table-format + concepts/egress-cost).
-
Separate the storage layer from the compute layer; use any Iceberg-compatible engine. The openness thesis is explicit: "you should own and be in control of your own data — separating the storage layer from the compute layer and allowing you to use the right query engine for the job." With Basin you "can read and write your data using any Iceberg-compatible engine, including PyIceberg, DuckDB, Snowflake, and Apache Spark" — and "that kind of data portability is only possible with free egress" (concepts/compute-storage-separation).
-
Basin Catalog auto-maintains tables for speed as datasets scale. "As datasets grow, Basin Catalog compacts metadata and data files to reduce I/O and generates statistics for query planning. Basin SQL uses those statistics to split queries into smaller tasks and distribute them across Workers, keeping queries fast and consistent as datasets scale." The catalog's maintenance loop (compaction + statistics) is the input to the query engine's planning + distributed scatter-gather execution.
-
Basin Pipelines scaled to 3 GB/s per stream. Since beta, Pipelines added: Cloudflare Logpush integration (transform Cloudflare logs with SQL → compressed Parquet or Iceberg tables), schema-aware Worker bindings (
wrangler typesgenerates TypeScript types from a stream's schema, catching missing fields / type mismatches before deployment), data-quality error visibility (the dashboard + GraphQL API surface dropped events, distinguishing missing fields / type mismatches / parse failures / null values), and full infrastructure-as-code (Terraform resources cover catalog, stream, sink, and the connecting SQL). Common pattern: transform Cloudflare HTTP logs during ingestion (sha256(ClientIP),WHERE EdgeResponseStatus >= 400) to reduce storage footprint and strip sensitive values before they are written. -
Basin SQL grew from a filter-and-explore engine into a full analytical query engine. At beta it was good at filtering/exploring large event + time-series tables; over the year it added 190+ scalar and aggregate functions (strings, timestamps, regex, cryptography, statistics, arrays, maps, structs), standard + approximate aggregations (
approx_distinct),GROUP BY/HAVING, inner/outer/semi/anti joins, subqueries, self-joins,DISTINCT/UNION/INTERSECT/EXCEPT, window functions,QUALIFY, grouping sets / rollups / cubes, CTEs,CASE, casting,EXPLAIN, and a JSON-function suite — queryable from Wrangler, the API, or a built-in dashboard editor (syntax highlighting, autocomplete, namespace/table browser, query stats + plans, exportable results). -
Serverless + usage-based pricing is the cost-efficiency lever. "You are only billed when Basin ingests, processes, or queries your data." No hourly charges, no separate infrastructure cost; hobby projects run "at little to no cost" while pricing "scales economically for larger enterprise use cases" (concepts/serverless-compute).
-
Existing configs keep working. "Existing Cloudflare Pipelines, R2 Data Catalog, and R2 SQL configurations will continue to work" — the GA rename is additive, not a breaking migration.
Operational numbers¶
- Basin Pipelines: up to 3 GB/s per stream ingest (expanded from beta); "tens of thousands of Pipelines" created since beta.
- Basin Catalog: "thousands of developers" since launch; created with
one command (
npx wrangler basin catalog create CATALOG_NAME); provides a "fully managed Apache Iceberg REST catalog." - Basin SQL: "hundreds of functions", specifically >190 scalar and aggregate functions.
- Adopters named: Anomaly (replaced a "complex AWS S3 and Athena setup" — Dax Raad, Co-Founder), Bobsled ("any region of every major data and AI platform… a fraction of the cost thanks to zero egress fees" — Julien Grobbelaar, Head of Platform); internal Cloudflare billing + infrastructure teams were among the first beta adopters.
- Fun framing: "roughly 20% of Earth's land drains into endorheic basins, much like over 20% of the web sits behind Cloudflare's network."
Basin Catalog maintenance features (added since beta)¶
- Per-table compaction policies — choose target file sizes based on each table's access pattern.
- Automatic snapshot expiration — remove old Iceberg snapshots per a retention policy while preserving a minimum number of recent snapshots.
- Unreferenced data-file cleanup — reclaim storage when snapshots expire, "without requiring a separate Spark maintenance job" (the OTF externalisation cost absorbed by the managed service).
- Manifest optimization — consolidate + cluster fragmented manifests by partition before compaction, reducing metadata I/O during query planning.
Roadmap (stated)¶
- Pipelines: custom partitioning when writing to Basin Catalog; schema migrations + updatable config/SQL; Iceberg V3 support incl. the Variant type for semi-structured data; stateful processing (streaming aggregations, joins, incrementally-updated materialized views).
- Basin Catalog: compaction that sorts/clusters data for query performance; more granular auth controls for namespaces/tables; jurisdiction support for data sovereignty.
- Basin SQL: advanced statistics + adaptive scheduling; full DDL support from Basin SQL; Iceberg V3 incl. VARIANT + geospatial types.
- Platform: latest Iceberg spec across the platform; push-button
ingestion sources/destinations with zero-config connections across
Cloudflare developer + observability products; adaptive table-maintenance
that organizes data around real query patterns; more real-time continuous
processing + action triggers; data-compliance / sovereignty tooling.
The north-star framing: "developers start with questions rather than
CREATEstatements or CLI commands" — data infrastructure "completely abstracted away."
Caveats¶
- GA/rebrand announcement — no independent benchmarks, no concrete query-latency or cost numbers beyond the 3 GB/s ingest figure and the function count.
- Roadmap items (Iceberg V3, Variant type, stateful processing, DDL, geospatial) are stated plans, not shipped.
- Basin SQL is described as "designed for reading large datasets" — DDL from Basin SQL is explicitly still future work, so writes still go through Pipelines or an external Iceberg engine today.
Systems / concepts extracted¶
- Systems: Basin (umbrella platform), Basin Pipelines, Basin Catalog, Basin SQL, Apache Iceberg, R2, K2, R2 Data Catalog (prior name), Cloudflare Logpush, DuckDB, Snowflake, Apache Parquet, Terraform.
- Concepts: concepts/compute-storage-separation, concepts/open-table-format, concepts/data-lakehouse, concepts/serverless-compute, concepts/egress-cost, concepts/scatter-gather-query, concepts/columnar-storage-format, concepts/elt-vs-etl, concepts/oltp-vs-olap.
Source¶
- Original: https://blog.cloudflare.com/cloudflare-basin/
- Raw markdown:
raw/cloudflare/2026-10-01-introducing-cloudflare-basin-an-open-serverless-data-platfor-9ee6a049.md
Related¶
- systems/basin — the GA-renamed umbrella platform.
- systems/basin-pipelines — ingestion / stream-processing component.
- systems/basin-catalog — managed Iceberg catalog component.
- systems/basin-sql — serverless distributed SQL engine component.
- systems/apache-iceberg — the open table format Basin is built on.
- systems/cloudflare-r2 — the zero-egress object-storage substrate.
- concepts/compute-storage-separation — the structural thesis.
- concepts/open-table-format — the portability thesis.
- concepts/data-lakehouse — the architectural class.
- companies/cloudflare — operator.