Skip to content

REDPANDA

Read original ↗

Multi-region high availability for Kafka workloads with a single Stretch Cluster

Summary

A Redpanda Operator v26.2 release post that hardens Redpanda on Kubernetes along two axes: the deployment architectures you can run safely, and the day-2 maintenance you perform on them. The headline is that Stretch Clusters go GA on Kubernetes — a single logical Redpanda cluster whose brokers span multiple K8s clusters (one per AZ / region / cloud), giving synchronous Raft-quorum multi-region replication with RPO = 0 and RTO = 0. This directly lifts the prior K8s substrate limitation (the 2025-02-11 stretch-cluster post said Self-Managed-on-K8s supported only multi-AZ). The release also ships per-broker pre/post-restart probes that gate rolling restarts on the cluster's actual replication state instead of a coarse cluster-wide health check, Gateway API support (TLSRoute for Kafka by SNI, HTTPRoute for Console), and a beta Pipeline CRD for Redpanda Connect.

Key takeaways

  1. Stretch Clusters are GA on Kubernetes (the substrate change). A Redpanda Stretch Cluster is a single logical cluster whose brokers span multiple Kubernetes clusters — typically one per AZ, region, or cloud provider. Because Redpanda replicates each partition with Raft and only acknowledges a write once a majority of replicas hold it durably, the topology gives "RPO = 0 … every acknowledged write already lives in a quorum that spans failure domains. Lose a region, and you lose zero acknowledged data" and "RTO = 0 … there's no failover step to perform. The surviving regions already hold a majority, so leadership re-elects automatically." Existing Kafka clients work unchanged through the Kafka-compatible API. (Source: this post)

  2. Two new CRDs model the stretch topology. The operator introduces StretchCluster and RedpandaBrokerPool in the cluster.redpanda.com/v1alpha2 API group. StretchCluster is the cluster-wide spec, created once — shared defaults for image, storage, resources, cluster config, and cross-cluster networking mode. RedpandaBrokerPool is created once per member Kubernetes cluster (per region/zone/cloud), references the parent StretchCluster via clusterRef, and carries region-specific overrides: replica count, external access, TLS, scheduling, storage. The operator stitches the pools into one cluster and keeps a consistent broker list across all member clusters. (Storage overrides must be set at creation time — volumeClaimTemplates on the rendered StatefulSet are immutable.)

  3. Contradiction with the 2025-02-11 post — resolved by this release. The 2025-02-11 stretch-cluster post stated "Self-Managed on K8s currently supports only multi-AZ deployments in all the cloud providers", so multi-region stretch previously required VMs / bare metal / cloud compute / Redpanda Cloud Dedicated + BYOC. Operator 26.2 removes that gap: multi-region stretch is now a first-class, GA Kubernetes deployment shape. (See the ## Contradiction sections on the stretch-cluster concept and Raft-quorum pattern pages.)

  4. Per-broker restart probes replace cluster-wide health gating (the day-2 change). Previously the operator gated restarts on a coarse cluster-wide health check, which left two windows open: before — the cluster could look healthy overall while one broker held the only in-sync replica for some partition, so restarting it could lose acks=1 data; after — a pod could report "Ready" before the broker had caught up its replicas, so rolling the next pod could drop the cluster under quorum. v26.2 sources two probes from Redpanda Core: pre-restart probe (/v1/broker/pre_restart_probe) asks Core whether this specific broker is safe to restart, checking three dangerous conditions — acks=1 data-loss risk, produce/consume unavailability from leaderless partitions, and acks=-1 produce rejection from losing quorum (RF=1 partitions are treated as acceptable, no redundancy by design); post-restart probe (/v1/broker/post_restart_probe) waits until the broker reports it has reclaimed its in-sync replicas before proceeding to the next pod (default 100% recovery, tunable via --post-restart-caught-up-percent).

  5. The probes fail closed and are automatic. Nothing to configure; the mechanism gracefully falls back to the old cluster-wide health check on Redpanda versions that don't expose the probes. Critically, the operator fails closed: "if a probe errors unexpectedly, the roll is deferred rather than barreling ahead." Net effect: upgrades and restarts respect the cluster's actual replication state, one broker at a time.

  6. Gateway API replaces Ingress for external access. Ingress was never a good fit for the Kafka protocol (HTTP-centric, and Kafka clients need to reconnect to specific brokers by hostname). v26.2 adds Gateway API support on two fronts: Console over HTTPRoute (GA) — declare a gateway stanza on the Console CR (or Helm values) and the operator renders + reconciles the HTTPRoute; Kafka over TLSRoute (beta) — the chart creates a bootstrap TLSRoute plus one per-broker TLSRoute, routed by SNI through a Gateway you manage (the broker advertises the same per-broker hostname the client reconnects on, so SNI matches what the Gateway routes by). Opt-in per listener enables gradual migration (move Kafka to TLSRoute while Admin / Schema Registry stay on NodePort/LoadBalancer). Gateway and Ingress listeners are mutually exclusive — enabling both fails fast.

  7. Redpanda Connect pipelines as first-class K8s resources (beta). A new Pipeline CRD lets you declare Connect pipelines the same way you manage clusters and topics: bind to a cluster via cluster.clusterRef and the operator injects broker addresses, TLS, and SASL creds into the redpanda input/output plugins (${RPK_BROKERS}, ${RPK_TLS_*} variables) — no hardcoded connection details. Supports userRef for SASL, per-pipeline serviceAccountName for cloud IAM (IRSA / Workload Identity / Pod Identity), valueSources for Secret/ConfigMap projection, paused: true to scale to zero without deleting, and standard K8s scheduling controls (nodeSelector, tolerations, topologySpreadConstraints, zones, disruption budget) so each pipeline lands on a node pool matched to its resource shape. Status conditions (ConfigValid, ClusterRef, Ready) report lifecycle position.

Operational specifics / numbers

  • API group: cluster.redpanda.com/v1alpha2 for StretchCluster, RedpandaBrokerPool, Pipeline, Console.
  • RPO = 0, RTO = 0 across regions/clouds for the stretch topology (the load-bearing claim; property of Raft quorum spanning failure domains).
  • Pre-restart probe checks: acks=1 data-loss risk; leaderless-partition produce/consume unavailability; acks=-1 produce rejection from quorum loss. RF=1 partitions accepted (no redundancy by design).
  • Post-restart probe default: wait for 100% in-sync-replica recovery; tunable via --post-restart-caught-up-percent for teams with recovery-time SLAs.
  • Gateway API CRDs are a prerequisite — not bundled; install standard-install.yaml (post cites v1.5.1) plus a Gateway controller (Envoy Gateway, Istio, Cilium, or NGINX Gateway Fabric). The chart deliberately does not create the Gateway — the infra team owns it; you reference it via parentRefs.
  • Kafka TLSRoute advertised port example: 9094; the Gateway does the listening.
  • Licensing: Stretch Clusters and Redpanda Connect pipelines are Enterprise features; Gateway API integration and the restart probes are available to everyone.

Caveats

  • Product-release / launch post (Tier-3 source). Included per AGENTS.md borderline rule: the architecture content (RPO/RTO=0 quorum semantics, replication-state-aware restart probes, SNI-based Kafka routing) is well over 20% of the body and is reusable distributed-systems design, not just feature PR.
  • Beta surfaces. Kafka TLSRoute, the Pipeline CRD, and the StretchCluster/RedpandaBrokerPool CRDs' v1alpha2 API may change before GA. Only Console-over-HTTPRoute and the restart probes are called out as broadly stable.
  • No quantitative benchmarks in this post (no measured failover times, cross-region RTT numbers, or throughput). RPO/RTO=0 is stated as a topology property, not measured here — the deeper stretch-cluster hazard analysis (cross-region RTT per write, bandwidth cost) lives in the 2025-02-11 source.
  • Cross-cluster networking mode (how member K8s clusters reach each other for Raft RPC) is named as a StretchCluster field but its mechanics are not detailed in this post.

Source

Last updated · 766 distilled / 2,225 read