Skip to content

SYSTEM Cited by 2 sources

Basin Pipelines

Basin Pipelines (a.k.a. Pipelines; formerly Cloudflare Pipelines) is Cloudflare's serverless ingestion / stream-processing service on the Developer Platform and the ingestion component of Basin: it accepts events through HTTP endpoints, Workers bindings, or Cloudflare Logpush; transforms them with a SQL query; and writes them to R2 (as JSON or Parquet files) or to a Basin Catalog (Apache Iceberg tables). Documented at developers.cloudflare.com/basin-pipelines.

GA capabilities (2026-10-01)

Renamed from Cloudflare Pipelines and GA'd on 2026-10-01. "Tens of thousands of Pipelines" created since beta. Since the beta launch:

  • Scaled to 3 GB/s per stream ingest.
  • Cloudflare Logpush integration — transform Cloudflare logs with SQL and store them as compressed Parquet files or Iceberg tables, ready to query with Basin SQL or another engine.
  • Schema-aware Worker bindings — running wrangler types generates TypeScript types from a stream's schema, "catching missing fields and type mismatches before deployment" (schema enforced at the binding boundary).
  • Data-quality error visibility — the dashboard + GraphQL API surface dropped events, distinguishing missing fields, type mismatches, parse failures, and null values.
  • Full infrastructure-as-code — Terraform resources cover the catalog, stream, sink, and the SQL that connects them.

A common ingestion-time transform pattern (reduce storage footprint, strip sensitive values before write):

INSERT INTO http_logs_sink
SELECT EdgeResponseStatus,
       to_timestamp_micros(EdgeStartTimestamp) AS event_time,
       upper(ClientRequestMethod) AS method,
       sha256(ClientIP) AS hashed_ip
FROM http_logs_stream
WHERE EdgeResponseStatus >= 400;

Roadmap (stated): custom partitioning when writing to Basin Catalog; schema migrations + updatable config/Pipelines SQL; Iceberg V3 incl. the Variant type for semi-structured data; stateful processing (streaming aggregations, joins, incrementally-updated materialized views — concepts/stateful-stream-processing).

Engine

Pipelines is powered by a pull-based stream-processing engine (Arroyo), which means "some other system has to store events before they are read, transformed, and written to R2." Because Cloudflare "commit[s] to never dropping events once they're accepted into the Pipelines Stream," that upstream store has to be durable over potentially long periods — which is the explicit motivation for building K2 as the durable buffer / ingestion layer underneath it.

Relationship to K2

  • K2 is the durable buffer; Pipelines is the pull-based processor. The Arroyo engine pulls events that K2 has durably stored, transforms them, and writes the results to R2 or Iceberg.
  • Choose Pipelines over K2 when "the end result is writing your events to object storage or Iceberg tables"; choose K2 "when doing custom processing or writing to other destinations."

Seen in

Last updated · 766 distilled / 2,225 read