Skip to content

SYSTEM Cited by 2 sources

Redpanda Operator

Redpanda Operator is Redpanda's Kubernetes Operator — the production-grade deployment path for Redpanda clusters on Kubernetes, positioned as the default recommendation over the Redpanda Helm chart. The operator manages cluster lifecycle via CRDs and provides five capabilities Helm cannot:

  1. Managed upgrade and rollback — safe rolling upgrades with reconciliation, vs Helm's manual-intervention rollback.
  2. Dynamic configuration — real-time config changes via CRDs, vs Helm's redeploy-to-apply-values flow.
  3. Advanced health checks and metrics — application-specific signals on top of generic Kubernetes probes.
  4. Lifecycle automation — scaling, failover, resource reconciliation, cleanup.
  5. Multi-tenancy management — one declarative CRD instance per cluster, vs separate Helm releases.

Evolution

The operator has gone through three phases ending in a 2025 consolidation:

Phase 1 (pre-2024): Two separate operators

Redpanda historically maintained two Kubernetes operators:

  • An internal Cloud operator managing Redpanda Cloud deployments.
  • A customer-facing operator for Self-Managed deployments.

The two were forked because the internal team and customers had divergent requirements. Over time, duplicated reconciliation logic became a cost.

Phase 2: Customer operator bundled FluxCD

To accelerate the customer operator's development, Redpanda bundled FluxCD — a GitOps controller similar to ArgoCD — and had the operator wrap the Redpanda Helm chart internally via FluxCD. This shipped the operator faster and kept operator-based and Helm-based deployments close.

The bundling turned into a anti-pattern:

  • Made the operator diverge from Kubernetes ecosystem norms.
  • Conflicted with customers already running their own FluxCD.
  • Coupled operator to bundled-FluxCD version and Helm-chart release.
  • Made merging the internal and customer operators harder.

Phase 3 (2025-): Unification + FluxCD removal

Three branches rolled the unification across 2025:

Branch FluxCD Redpanda core dep Notes
v2.3.x optional (spec.chartRef.useFlux) Helm chart Opt-out available
v2.4.x disabled by default (Jan 2025) Helm chart Same toggle
v25.1.x removed removed Version-aligned, beta

The v2.4.x → v25.1.x version jump is deliberate — the operator adopts a version-aligned compatibility scheme where the operator/chart version number matches the Redpanda core version, retiring the prior compatibility matrix. Each operator/chart version is compatible with the Redpanda version immediately above and below it (±1 minor window).

(Source: sources/2025-05-06-redpanda-a-guide-to-redpanda-on-kubernetes)

Redpanda Operator v26.2 (2026-08-11)

The v26.2 release hardens Redpanda on Kubernetes along two axes at once — the deployment architectures you can run safely, and the day-2 maintenance you perform on them. (Source: sources/2026-08-11-redpanda-multi-region-high-availability-for-kafka-workloads-with-a-single-stretch-cluster)

Stretch Clusters GA on Kubernetes — two new CRDs

The headline: Stretch Clusters go GA on Kubernetes, closing the substrate gap noted below. A single logical Redpanda cluster now spans multiple Kubernetes clusters (typically one per AZ / region / cloud), giving synchronous Raft-quorum multi-region replication with RPO = 0 and RTO = 0 — existing Kafka clients work unchanged. Two new CRDs in cluster.redpanda.com/v1alpha2 model it:

  • StretchCluster — the cluster-wide spec, created once. Holds shared defaults: image, storage, resources, cluster config, and cross-cluster networking mode.
  • RedpandaBrokerPool — created once per member Kubernetes cluster (per region/zone/cloud). References the parent StretchCluster via clusterRef and carries region-specific overrides: replica count, external access, TLS, scheduling, storage. (Storage overrides must be set at creation time — volumeClaimTemplates on the rendered StatefulSet are immutable.)

The operator stitches the pools into one cluster and keeps a consistent broker list across all member clusters. Stretch Clusters are an Enterprise feature.

Replication-aware rolling restarts — per-broker probes

v26.2 replaces the old coarse cluster-wide health gate with per-broker restart safety probes sourced from Redpanda Core, implementing a replication-aware rolling restart:

  • Pre-restart probe (/v1/broker/pre_restart_probe) — before touching a broker, Core checks its actual partition leadership and replication state for three dangerous conditions: acks=1 data-loss risk, produce/consume unavailability from leaderless partitions, and acks=-1 produce rejection from losing quorum. RF=1 partitions are treated as acceptable.
  • Post-restart probe (/v1/broker/post_restart_probe) — after a broker returns, waits until it has reclaimed its in-sync replicas before proceeding to the next pod. Default 100% recovery, tunable via --post-restart-caught-up-percent.

Automatic (nothing to configure), fails closed (a probe error defers the roll), and gracefully falls back to the old cluster-wide health check on Redpanda versions that don't expose the probes. Available to everyone.

Gateway API integration

Adds Kubernetes Gateway API support for external access, a better fit than Ingress for the Kafka protocol:

  • Console over HTTPRoute (GA) — declare a gateway stanza on the Console CR (or Helm values); the operator renders and reconciles the HTTPRoute against an infra-team-owned Gateway.
  • Kafka over TLSRoute (beta) — a bootstrap TLSRoute plus one per-broker TLSRoute, routed by SNI through a Gateway you manage. Opt-in per listener (migrate Kafka while Admin / Schema Registry stay on NodePort/LoadBalancer). Gateway and Ingress are mutually exclusive per listener. See gateway-api-route-attachment. Available to everyone.

Redpanda Connect pipelines as K8s resources (beta)

A new Pipeline CRD lets you declare Redpanda Connect pipelines declaratively: bind to a cluster via cluster.clusterRef and the operator injects broker addresses, TLS, and SASL creds (${RPK_BROKERS}, ${RPK_TLS_*}). Supports userRef (SASL), per-pipeline serviceAccountName (IRSA / Workload Identity / Pod Identity), valueSources (Secret/ConfigMap projection), paused: true (scale to zero without deleting), full K8s scheduling controls, and ConfigValid/ClusterRef/Ready status conditions. Enterprise feature.

Deployment-shape limitation

Per the 2025-02-11 stretch-cluster post, "Self-Managed on K8s currently supports only multi-AZ deployments in all the cloud providers" — at that time multi-region stretch clusters were not supported on Kubernetes; they required VMs, bare metal, cloud compute, or Redpanda Cloud Dedicated / BYOC. The 2025-05-06 deployment guide did not revisit this substrate constraint.

This limitation was removed in Operator v26.2 (2026-08-11): Stretch Clusters are now GA on Kubernetes (see the v26.2 section above and the ## Contradiction on multi-region-stretch-cluster).

Last updated · 766 distilled / 2,225 read