Skip to content

AWS 2026-09-01

Read original ↗

Hybrid cloud orchestration: Modernizing on-premises infrastructure management with AWS

Part 1 of an AWS Architecture Blog series on building a hybrid cloud orchestration solution: a centralized, event-driven orchestration engine running on AWS serverless services that manages geographically dispersed on-premises data centers with thousands of bare-metal servers — hardware lifecycle, OS, and Amazon EKS Anywhere Kubernetes clusters — while keeping execution on-premises. The design target is environments where the Kubernetes control plane must stay on-prem for data-sovereignty, regulatory, or DDIL (disconnected, disrupted, intermittent, or limited network) reasons. It is the canonical wiki instance of centralized control plane on AWS + distributed execution on-premises, glued together by an event-driven orchestration engine.

Summary

The solution has three layers: (1) a centralized orchestration engine on AWS built from API Gateway + Lambda + Step Functions + EventBridge + a DynamoDB-backed Inventory Management System as the single source of truth; (2) distributed on-premises infrastructure running EKS Anywhere clusters on bare-metal servers managed through vendor-agnostic Redfish APIs; and (3) hybrid connectivity (Direct Connect or Site-to-Site VPN) linking the two. Operators issue Orders — trackable lifecycle operations that emit events, get routed by EventBridge to the right Step Functions workflow, and update status asynchronously in DynamoDB. Because on-prem operations (firmware updates, cluster builds) can run for hours, the engine leans on the Step Functions callback (task token) pattern and the Distributed Map state to fan out to thousands of resources. Observability is centralized by shipping on-prem telemetry (ADOT → Managed Prometheus → Managed Grafana) back to the AWS Region.

Key takeaways

  1. Centralized orchestration, distributed execution. AWS holds the control plane (API, workflows, state); on-prem sites run the actual hardware and Kubernetes operations. This is control-plane / data-plane separation applied across the cloud/on-prem boundary — and the reason it works under DDIL conditions where the network to AWS may be unreliable. (Source: sources/2026-09-01-aws-hybrid-cloud-orchestration-modernizing-on-premises-infrastructure)
  2. The Inventory Management System is the single source of truth. DynamoDB tables track sites, hardware (BIOS/firmware versions, encrypted credentials, IP/MAC, rack position, cluster membership), clusters (K8s config, node groups, addon versions, management↔workload relationships), orders (operation lifecycle), and a catalog of reusable blueprints/templates. It records both current state and operational history / resource relationships.
  3. Orders decouple request from execution. An API call like /clusters/{id}/terminate creates a DynamoDB record, returns an order ID immediately, and runs the workflow asynchronously — operations that take minutes to hours don't block the caller. AWS-managed Step Functions events + custom workflow events progressively update order status.
  4. The callback pattern bridges to slow on-prem systems. Step Functions can pause, hand a task token to an external system, and resume only when that system calls back — essential for firmware updates or cluster deployments that run for hours on-premises. (Source: sources/2026-09-01-aws-hybrid-cloud-orchestration-modernizing-on-premises-infrastructure)
  5. Distributed Map scales one-server operations to thousands. A workflow that manages the power state of a single server scales to manage power across thousands of servers simultaneously across sites via the Step Functions Distributed Map state.
  6. Redfish gives vendor-agnostic hardware control. Redfish (a DMTF standard protocol) delivers standardized APIs for BIOS configuration, firmware/NIC updates, power management, and health monitoring across heterogeneous bare-metal hardware — the vendor-agnostic hardware API that makes fleet-wide hardware automation possible.
  7. EKS Anywhere runs the whole cluster on-prem. EKS Anywhere creates/operates Kubernetes clusters on customer hardware using the same Amazon EKS Distro, via the bare-metal provider (config file + hardware CSV → network boot → OS + K8s install → cluster up). Management clusters host orchestration; workload clusters host applications; the mapping is tracked in inventory. EKS commands run against on-prem servers via Systems Manager (SSM) + Batch. EKS Hybrid Nodes is the recommended alternative when the K8s control plane can live in the cloud and connectivity is reliable.
  8. Conflict management via the inventory lock. New orders are denied if another operation is running on the same resource (conflict-prevention-via-inventory-lock) — preventing scenarios like scaling a cluster during an upgrade.
  9. Extensibility without touching core logic. New resource types / operations are added by implementing a Step Functions workflow and registering an EventBridge rule mapping the API endpoint to it; the order-tracking core stays unchanged.
  10. Centralized observability closes the fragmented-visibility gap. AWS Distro for OpenTelemetry (ADOT) collectors on each EKS Anywhere cluster scrape server/K8s/app metrics and forward to Amazon Managed Service for Prometheus; Amazon Managed Grafana provides unified dashboards + alerting across the whole distributed estate regardless of hardware vendor.

Architecture detail

Foundational objects

  • Site — physical location or logical grouping (central / regional / edge data center); carries location-specific policies, connectivity, and compliance controls.
  • Server — bare-metal server providing compute/storage/network for containers; hardware managed via Redfish.
  • Cluster — an EKS Anywhere cluster; management clusters run orchestration components, workload clusters host applications.
  • Order — a trackable lifecycle operation executed as a workflow, with a unique ID; emits an event → EventBridge routes to the corresponding Step Functions workflow → status updated from workflow state-change events.

Security and configuration

  • SSM Parameter Store for centralized config; Secrets Manager for credentials.
  • IAM roles for fine-grained access, with IAM Roles Anywhere extending AWS access to on-premises clusters without long-term credentials — the orchestration engine registers each cluster's own CA certificate as its trust anchor at bring-up.
  • SSM hybrid activations register on-prem instances so the SSM agent can manage them alongside cloud resources.
  • AWS Private CA manages certificates for secure orchestrator↔on-prem communication; every component runs least-privilege.

Hardware and cluster lifecycle

  • Hardware ops via Redfish: firmware management, NIC upgrades, power (reboot/shutdown/cycle), health (CPU/memory/disk), BIOS golden templates.
  • Cluster ops via EKS Anywhere: creation, scaling (add/remove workers), coordinated version upgrades, graceful termination. Cluster creation is a sequenced flow (hardware selection by placement strategy → pre-flight → bootstrap admin machine → on-prem command execution + await → add-ons → post-deployment health checks), orchestrated with Step Functions child workflows, callback patterns, and dependency mapping.

Hybrid integration patterns

Even though clusters run on-prem, workloads depend on cloud-spanning capabilities, applied automatically from inventory as state changes:

  • Automated DNS — DynamoDB Streams trigger Lambda to update Route 53 private hosted zones; Route 53 Resolver endpoints make records resolvable from both AWS and on-prem (dynamodb-streams-triggered-dependent-infra).
  • Certificate lifecycle — AWS Private CA as managed CA; cert-manager + the AWS Private CA Issuer request/renew/distribute certs automatically (no per-site CA).
  • Secure AWS access — IAM Roles Anywhere issues short-lived credentials in exchange for a certificate the workload holds; accepted only if it chains to a trusted anchor (the cluster's own registered CA).
  • Persistent storage — external storage such as Portworx integrated via DynamoDB-Streams-triggered storage-provider API calls during node provisioning/cleanup.

Operational numbers / caveats

  • Genre: Part-1 reference architecture, not a production retrospective — no throughput, latency, or cost numbers are disclosed; subsequent posts promise code examples, deployment templates, and detailed workflows.
  • Scale framing is qualitative: "thousands of servers," "hundreds of sites."
  • EKS Anywhere places full cluster lifecycle responsibility on the customer — the orchestration engine exists precisely to automate that work across sites. EKS Hybrid Nodes is the recommended alternative when the control plane can be cloud-managed.

Source

Last updated · 766 distilled / 2,225 read