Announcing Cloudflare K2: serverless event streams¶
Summary¶
Cloudflare launched K2 (public beta, 2026-10-01), a durable event-streaming primitive on the Developer Platform: producers send events to a K2 stream, which stores them as an ordered log, and consumers read them back either by splitting reads across a set of consumers (read parallelism) or by delivering all messages to all consumers (pub/sub fan-out). The architectural bet is that K2 implements a partitioned, durable log on top of R2 object storage rather than deploying Kafka — because Cloudflare's edge (335+ cities) gets small, ephemeral slices of machines over the public Internet and "often cannot run traditional distributed systems software like Kafka." Offloading replication and consensus to R2's 11-nines durable, strongly-consistent storage lets K2 keep the compute layer "radically simpler, cheaper, and higher performance" and scale compute and storage independently. K2 was first built as the durable buffer behind Basin Pipelines' pull-based stream-processing engine.
Key takeaways¶
-
Durable log on object storage, not on local disk or a Kafka cluster. K2 is a partitioned durable log built on R2. Replication and consensus are "offloaded to the storage layer", so the K2 application tier carries no quorum/consensus of its own — the canonical object-storage-as-disk-root / compute-storage separation move applied to a streaming broker. "Offloading replication and consensus to the storage layer allows us to make the application layer (K2 in this case) radically simpler, cheaper, and higher performance." (Source)
-
Edge constraints forced the design. "Pipelines runs on the Cloudflare edge, which spans a huge number of servers across over 335 cities. Our unique architecture means we often cannot run traditional distributed systems software like Kafka." Stateful services at the edge get "relatively small slices of machines, those machines are relatively ephemeral, and networking is often over the public Internet" — the opposite of Kafka's assumed dedicated, long-lived, LAN-connected brokers.
-
No append on object stores → accumulate-then-write segments. R2, like other object stores, "does not support appends, the standard operation on a log." K2 instead "accumulat[es] writes in-memory on an edge service. After waiting a short period for data to arrive, we write all events as a segment file" — the micro-batch / small-file-avoidance discipline: segments must be "large enough to overcome the cost of writing and reading each one."
-
Ordering + offsets come from R2's atomic operations, not a coordinator. "We achieve ordering and strictly incrementing offsets using R2's atomic operations without needing a separate coordination service" — a conditional/atomic-write on the storage layer substitutes for a leader-election / consensus service (strictly-incrementing offsets).
-
The cost is produce latency. "Writing to object storage is slower than a local disk, and we have to wait for the local batch to accumulate before starting the write. In our initial release of K2, this adds up to about 1 second of produce latency at the 99th percentile." The explicit batching-latency tradeoff: object-store durability + batch-before-write buys cost and simplicity at the price of p99 produce latency.
-
Two consumption shapes: read-parallel subscriptions and pub/sub fan-out. A subscription divides work across consumers (read parallelism — "scaling out to multiple readers to handle more load than a single server can manage"); a separate subscription per consumer gives every consumer all messages (pub/sub). "Or you can mix-and-match … having multiple independent consumer pools." This is the decoupling that solves the producer/consumer scale-and-time mismatch the post opens with.
-
Lease / ack / nack / extend delivery contract. On
consume, a client receives a lease for a batch of events (5 minutes). It can ack (mark processed, never redelivered), nack (processing failed, redeliver), or extend the lease (needs more time). Lease-expiry-redelivers makes K2 an at-least-once system — duplicates are possible (crash after processing but before ack), so consumers must be idempotent. -
Batch-grained, not message-grained. "Messages are produced and consumed as batches — enabling efficient processing at the expense of message-level retries." K2 deliberately gives up the per-item retry/delay/DLQ machinery that Queues offers, in exchange for high-scale data movement, long retention, and fan-out. The batching "also drives higher producer latency than for queues."
-
Positioned against Queues and Pipelines, not replacing them. Queues = individual expensive work items with per-item retries/delays/DLQ; K2 = high-scale data movement + long-term retention + fan-out; Basin Pipelines = when the end result is writing events to R2 or an Iceberg (Basin Catalog) table. "K2 when doing custom processing or writing to other destinations."
Systems / concepts / patterns extracted¶
- System (new): Cloudflare K2 — durable partitioned event-streaming log on R2.
- Systems (enriched): R2 (durable-log substrate — sixth substrate role), Kafka (the explicit "this is where most companies would deploy Kafka" counterfactual K2 avoids at the edge), Queues (sibling messaging primitive — contrast), Basin Pipelines (first K2 consumer / motivating workload), cf / Wrangler (stream-creation surfaces).
- Concepts (enriched, no new pages): replicated log, object storage as disk root, compute-storage separation, small-file problem on object storage, micro-batching, offset-preserving replication (strictly-incrementing offsets via atomic ops), at-least-once delivery, backpressure (the open-paragraph producer/consumer scale mismatch K2's buffer absorbs), batching-latency tradeoff.
- Patterns (enriched): tiered storage to object store (K2 is the limit case — object store is the only tier, not the cold tier), batch over network to broker (K2 batches on the edge service before the segment write), conditional write (R2 atomic ops for offset ordering).
Taxonomy gate¶
NO NEW CONCEPT OR PATTERN PAGES. Every architectural idea in the post maps INTO existing controlled vocabulary:
- log-on-object-storage → folds into concepts/object-storage-as-disk-root + concepts/compute-storage-separation + patterns/tiered-storage-to-object-store (K2 is the degenerate "object store is the whole log" case).
- accumulate-in-memory-then-write-segment → folds into concepts/micro-batching + concepts/small-file-problem-on-object-storage
- patterns/batch-over-network-to-broker.
- atomic-offset-via-object-store → folds into patterns/conditional-write + concepts/offset-preserving-replication.
- lease-ack-nack-extend delivery contract → recorded as prose on the K2 system page; it's K2-specific API mechanics, not a reusable named pattern (and it's single-source). Its semantic consequence (at-least-once + idempotent consumers) maps into concepts/at-least-once-delivery.
Only systems/cloudflare-k2 is minted — a proper-noun system (exempt from the concept/pattern page-creation bar).
Operational numbers¶
- 335+ cities span the Cloudflare edge K2 runs on.
- ~1 second produce latency at p99 in the initial release (object-store write + batch-accumulation wait).
- R2 durability: 11 nines (the storage substrate's guarantee K2 inherits).
- 5-minute consume lease per batch (extendable).
- Beta limits: max 10 GB storage, 30 MB/s produce per stream.
- Default stream
retention_seconds: 604800 (7 days) in the create example. - Anticipated post-beta pricing: $0.04/GB produced, $0.04/GB consumed, $0.02/GB/month retained; unbilled during beta.
- Available to Workers Paid accounts.
Roadmap (stated)¶
- Higher write parallelism — up to multi-GB/s streams.
- Message keys + key-based ordering guarantees.
- Push-based worker consumers.
- An Express tier with lower produce + end-to-end latency.
- Drop-in Apache Kafka client support.
Caveats¶
- Beta: hard caps (10 GB / 30 MB/s per stream), unbilled, limited guarantees (no key-based ordering yet; batch-grained only — no message-level retries).
- p99 produce latency ~1 s is a current-release number, explicitly flagged for improvement via the Express tier.
- No independent benchmarks; all numbers are Cloudflare-stated. A promised "upcoming technical deep dive" will carry the real design internals (segment format, R2 atomic-op mechanics, partitioning scheme) — this launch post is architecture-at-a-level, not the full design doc.
Source¶
- Original: https://blog.cloudflare.com/cloudflare-k2-streams/
- Raw markdown:
raw/cloudflare/2026-10-01-announcing-cloudflare-k2-serverless-event-streams-d7084788.md
Related¶
- systems/cloudflare-k2 — the system this source introduces.
- systems/cloudflare-r2 — the durable-log storage substrate.
- systems/kafka — the counterfactual K2 avoids running at the edge.
- systems/cloudflare-queues — sibling messaging primitive (contrast).
- concepts/object-storage-as-disk-root — the architectural root move.
- concepts/compute-storage-separation — independent scaling of each.
- concepts/replicated-log — the abstraction K2 implements.
- patterns/tiered-storage-to-object-store — K2 as the limit case.
- companies/cloudflare — operator.