Multi-region high availability for Kafka workloads with a single Stretch Cluster¶
Summary¶
A Redpanda Operator v26.2 release post that hardens Redpanda on Kubernetes along two axes: the deployment architectures you can run safely, and the day-2 maintenance you perform on them. The headline is that Stretch Clusters go GA on Kubernetes — a single logical Redpanda cluster whose brokers span multiple K8s clusters (one per AZ / region / cloud), giving synchronous Raft-quorum multi-region replication with RPO = 0 and RTO = 0. This directly lifts the prior K8s substrate limitation (the 2025-02-11 stretch-cluster post said Self-Managed-on-K8s supported only multi-AZ). The release also ships per-broker pre/post-restart probes that gate rolling restarts on the cluster's actual replication state instead of a coarse cluster-wide health check, Gateway API support (TLSRoute for Kafka by SNI, HTTPRoute for Console), and a beta Pipeline CRD for Redpanda Connect.
Key takeaways¶
-
Stretch Clusters are GA on Kubernetes (the substrate change). A Redpanda Stretch Cluster is a single logical cluster whose brokers span multiple Kubernetes clusters — typically one per AZ, region, or cloud provider. Because Redpanda replicates each partition with Raft and only acknowledges a write once a majority of replicas hold it durably, the topology gives "RPO = 0 … every acknowledged write already lives in a quorum that spans failure domains. Lose a region, and you lose zero acknowledged data" and "RTO = 0 … there's no failover step to perform. The surviving regions already hold a majority, so leadership re-elects automatically." Existing Kafka clients work unchanged through the Kafka-compatible API. (Source: this post)
-
Two new CRDs model the stretch topology. The operator introduces
StretchClusterandRedpandaBrokerPoolin thecluster.redpanda.com/v1alpha2API group. StretchCluster is the cluster-wide spec, created once — shared defaults for image, storage, resources, cluster config, and cross-cluster networking mode. RedpandaBrokerPool is created once per member Kubernetes cluster (per region/zone/cloud), references the parentStretchClusterviaclusterRef, and carries region-specific overrides: replica count, external access, TLS, scheduling, storage. The operator stitches the pools into one cluster and keeps a consistent broker list across all member clusters. (Storage overrides must be set at creation time —volumeClaimTemplateson the rendered StatefulSet are immutable.) -
Contradiction with the 2025-02-11 post — resolved by this release. The 2025-02-11 stretch-cluster post stated "Self-Managed on K8s currently supports only multi-AZ deployments in all the cloud providers", so multi-region stretch previously required VMs / bare metal / cloud compute / Redpanda Cloud Dedicated + BYOC. Operator 26.2 removes that gap: multi-region stretch is now a first-class, GA Kubernetes deployment shape. (See the
## Contradictionsections on the stretch-cluster concept and Raft-quorum pattern pages.) -
Per-broker restart probes replace cluster-wide health gating (the day-2 change). Previously the operator gated restarts on a coarse cluster-wide health check, which left two windows open: before — the cluster could look healthy overall while one broker held the only in-sync replica for some partition, so restarting it could lose
acks=1data; after — a pod could report "Ready" before the broker had caught up its replicas, so rolling the next pod could drop the cluster under quorum. v26.2 sources two probes from Redpanda Core: pre-restart probe (/v1/broker/pre_restart_probe) asks Core whether this specific broker is safe to restart, checking three dangerous conditions —acks=1data-loss risk, produce/consume unavailability from leaderless partitions, andacks=-1produce rejection from losing quorum (RF=1 partitions are treated as acceptable, no redundancy by design); post-restart probe (/v1/broker/post_restart_probe) waits until the broker reports it has reclaimed its in-sync replicas before proceeding to the next pod (default 100% recovery, tunable via--post-restart-caught-up-percent). -
The probes fail closed and are automatic. Nothing to configure; the mechanism gracefully falls back to the old cluster-wide health check on Redpanda versions that don't expose the probes. Critically, the operator fails closed: "if a probe errors unexpectedly, the roll is deferred rather than barreling ahead." Net effect: upgrades and restarts respect the cluster's actual replication state, one broker at a time.
-
Gateway API replaces Ingress for external access. Ingress was never a good fit for the Kafka protocol (HTTP-centric, and Kafka clients need to reconnect to specific brokers by hostname). v26.2 adds Gateway API support on two fronts: Console over HTTPRoute (GA) — declare a
gatewaystanza on the Console CR (or Helm values) and the operator renders + reconciles the HTTPRoute; Kafka over TLSRoute (beta) — the chart creates a bootstrap TLSRoute plus one per-broker TLSRoute, routed by SNI through a Gateway you manage (the broker advertises the same per-broker hostname the client reconnects on, so SNI matches what the Gateway routes by). Opt-in per listener enables gradual migration (move Kafka to TLSRoute while Admin / Schema Registry stay on NodePort/LoadBalancer). Gateway and Ingress listeners are mutually exclusive — enabling both fails fast. -
Redpanda Connect pipelines as first-class K8s resources (beta). A new
PipelineCRD lets you declare Connect pipelines the same way you manage clusters and topics: bind to a cluster viacluster.clusterRefand the operator injects broker addresses, TLS, and SASL creds into the redpanda input/output plugins (${RPK_BROKERS},${RPK_TLS_*}variables) — no hardcoded connection details. SupportsuserReffor SASL, per-pipelineserviceAccountNamefor cloud IAM (IRSA / Workload Identity / Pod Identity),valueSourcesfor Secret/ConfigMap projection,paused: trueto scale to zero without deleting, and standard K8s scheduling controls (nodeSelector,tolerations,topologySpreadConstraints, zones, disruption budget) so each pipeline lands on a node pool matched to its resource shape. Status conditions (ConfigValid,ClusterRef,Ready) report lifecycle position.
Operational specifics / numbers¶
- API group:
cluster.redpanda.com/v1alpha2forStretchCluster,RedpandaBrokerPool,Pipeline,Console. - RPO = 0, RTO = 0 across regions/clouds for the stretch topology (the load-bearing claim; property of Raft quorum spanning failure domains).
- Pre-restart probe checks:
acks=1data-loss risk; leaderless-partition produce/consume unavailability;acks=-1produce rejection from quorum loss. RF=1 partitions accepted (no redundancy by design). - Post-restart probe default: wait for 100% in-sync-replica
recovery; tunable via
--post-restart-caught-up-percentfor teams with recovery-time SLAs. - Gateway API CRDs are a prerequisite — not bundled; install
standard-install.yaml(post citesv1.5.1) plus a Gateway controller (Envoy Gateway, Istio, Cilium, or NGINX Gateway Fabric). The chart deliberately does not create the Gateway — the infra team owns it; you reference it viaparentRefs. - Kafka TLSRoute advertised port example:
9094; the Gateway does the listening. - Licensing: Stretch Clusters and Redpanda Connect pipelines are Enterprise features; Gateway API integration and the restart probes are available to everyone.
Caveats¶
- Product-release / launch post (Tier-3 source). Included per AGENTS.md borderline rule: the architecture content (RPO/RTO=0 quorum semantics, replication-state-aware restart probes, SNI-based Kafka routing) is well over 20% of the body and is reusable distributed-systems design, not just feature PR.
- Beta surfaces. Kafka TLSRoute, the Pipeline CRD, and the
StretchCluster/RedpandaBrokerPool CRDs'
v1alpha2API may change before GA. Only Console-over-HTTPRoute and the restart probes are called out as broadly stable. - No quantitative benchmarks in this post (no measured failover times, cross-region RTT numbers, or throughput). RPO/RTO=0 is stated as a topology property, not measured here — the deeper stretch-cluster hazard analysis (cross-region RTT per write, bandwidth cost) lives in the 2025-02-11 source.
- Cross-cluster networking mode (how member K8s clusters reach
each other for Raft RPC) is named as a
StretchClusterfield but its mechanics are not detailed in this post.
Source¶
- Original: https://www.redpanda.com/blog/multi-region-high-availability-kafka-stretch-clusters
- Raw markdown:
raw/redpanda/2026-08-11-multi-region-high-availability-for-kafka-workloads-with-a-si-c2aa35b7.md
Related¶
- systems/redpanda-operator — the subject; v26.2 features.
- systems/redpanda, systems/redpanda-connect, systems/gateway-api, systems/kubernetes
- multi-region-stretch-cluster — now GA on K8s.
- replication-aware-rolling-restart — the day-2 idea.
- concepts/rpo-rto, in-sync-replica-set, concepts/kubernetes-operator-pattern, concepts/fail-open-vs-fail-closed
- multi-region-raft-quorum — the RPO=0 shape.
- per-broker-restart-safety-probe — the new pattern.
- companies/redpanda