Skip to content

AWS 2026-10-02

Read original ↗

Deploy open source Regional availability tools in your VPC

Summary

AWS publishes Regional availability data — which services, features, API operations, and CloudFormation resource types exist in each Region — and already exposes it three ways (the Capabilities by Region explorer page, an S3 access point for pipeline integration, and the AWS Knowledge MCP server). This post addresses a fourth demand teams raised: run that data as infrastructure you own, inside your own VPC, refreshing on your schedule, filtered to your workload — i.e. under your own governance for data-residency and compliance reporting. It introduces two open-source solutions. Capability Insights for AWS deploys a searchable Regional-availability dashboard into the customer's own account; a scheduled Lambda (running outside the VPC) is the only component that leaves the account, pulling the published dataset from the Capabilities by Region S3 access point every 24h and writing it to an in-account website bucket. Workload Analysis is an additive stack that narrows the 200+ service / 35+ Region catalog to the 20–30 services the account actually runs, turning a multi-week Regional-expansion gap analysis into a ~30-minute review. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)

Key takeaways

  1. "Deploy the data as infrastructure you own" is the thesis. The published dataset is consumable from S3 or MCP, but teams wanted it inside their perimeter — their VPC, their account, their refresh cadence, their governance — so it runs "the same way you run the rest of your stack." Three levels of control: consume from S3, deploy into your VPC, or filter to your workload. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
  2. Minimized egress by moving one function outside the VPC. The data-fetch Lambda runs outside the VPC so it can read the AWS-published S3 access point over the AWS network, then writes into the in-account website bucket; the dashboard reads that bucket through an S3 gateway endpoint, and the private API is reached through an API Gateway VPC endpoint. "The only call that leaves your account" during normal operation is the scheduled dataset read. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
  3. EventBridge schedule drives a 24h refresh. An EventBridge schedule invokes the data-fetch Lambda every 24 hours; the dashboard "auto-refreshes every 24 hours," so there are no external dependencies during normal operation once seeded. The API Lambda can also invoke data-fetch on demand through the Lambda VPC endpoint. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
  4. Workload Analysis is a Step Functions state machine with two parallel analyzers + a merge step (a fan-out/fan-in orchestration):
  5. CloudTrail Analyzer — creates an AWS Glue Data Catalog table over the account's CloudTrail log bucket, then runs an Athena query (with Lake Formation catalog access) extracting distinct service/API/Region/account combinations from the last N days (default 30). This finds what the account calls.
  6. CloudFormation Analyzer — calls ListStacks + GetTemplate on every active stack, extracts AWS:: resource types and scalar property values, maps them to service names, and records which stack contributed each resource. This finds what the account deploys (not just calls).
  7. Usage Decorator — runs after both complete; reads the primary catalogs (products.json, apis.json, cfn_resources.json) from the website bucket, intersects with the analyzer outputs, and writes personalized files back for the dashboard/API to consume. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
  8. "Calls" vs "deploys" is a deliberate two-signal design. CloudTrail tells you which services the account invoked over a window; CloudFormation tells you which resource types the account provisions and even which property configurations are in use (e.g. AWS::EC2::Instance with InstanceType: t3.medium from the app stack and m5.xlarge from the data-processing stack), each traced to its source stack. The union is a truer picture of a workload's real surface than either signal alone. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
  9. The personalization is the point for gap analysis. Without the filter the dashboard shows 200+ services across 35 Regions; with the combined filter a "typical web application" account drops to 28 services with per-stack attribution. Planning expansion into, e.g., Europe (Zurich) eu-central-2, you scan a personalized list where "every row is something your stacks deploy" instead of evaluating the full catalog against a Region matrix. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
  10. Private-only access by design. The dashboard is an S3-hosted static website reachable only from within the VPC; access requires VPN / Direct Connect, an AWS Client VPN endpoint, or an EC2 SOCKS-proxy jump host. The API Gateway endpoint is private — API requests must originate inside the VPC. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
  11. Scoped runtime IAM. The CloudFormation-created roles use scoped permissions at runtime: S3 object read/write, CloudFormation read, Athena queries, Glue + Lake Formation catalog access, Lambda invocation, and Step Functions execution. The source account 686591367145 is the AWS-managed account hosting the public dataset's S3 access point (ARN used verbatim). (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)

Architecture

Capability Insights for AWS (Part 1): - EventBridge schedule → data-fetch Lambda (outside VPC) → reads arn:aws:s3:...:accesspoint/aws-capabilities-public → writes website bucket (in account). - Client in public subnet → dashboard via S3 gateway endpoint; private API via API Gateway VPC endpoint; API Lambda can invoke data-fetch via Lambda VPC endpoint. - Deployment-assets bucket supplies Lambda code during deploy only. - Website served at http://capability-insights-website-<ACCOUNT_ID>-<REGION>.s3-website-<REGION>.amazonaws.com (VPC-internal).

Workload Analysis (Part 2): - EventBridge schedule or on-demand POST /analysis → Step Functions state machine. - Parallel branch: CloudTrail Analyzer (Glue table + Athena over CloudTrail bucket) ∥ CloudFormation Analyzer (ListStacks/GetTemplate). - Join: Usage Decorator intersects analyzer outputs with primary catalogs, writes personalized files to the website bucket. - Deployed additively with --enable-usage-analysis --cloudtrail-bucket <bucket>; wires the API Lambda to trigger analysis and serve personalized results.

Operational numbers

  • Dashboard refresh: every 24 hours (EventBridge schedule).
  • CloudTrail scan window: last N days, default 30 (daysToScan).
  • Catalog scope: 200+ services across 35+ Regions; filtered to 20–30 (worked example: 28) services an account runs.
  • Gap-analysis effort: multi-week → ~30-minute review once filtered.
  • Workload Analysis first run: 2–5 minutes depending on account size / active stack count.
  • Prereqs: Node.js v24.18.0; VPC with DNS resolution + hostnames, one public subnet (IGW) + one private subnet (no internet route), both with an S3 gateway VPC endpoint; active CloudTrail landing in S3 (Part 2).
  • Cost: standard service pricing for Lambda / S3 / API Gateway / Athena / Step Functions; no additional charge for the solution itself.

Caveats

  • Reference/how-to post, not a production retrospective. No named customer, no production incident; the "multi-week → 30-minute" and "28 services" figures are an illustrative worked example, not measured fleet data.
  • Open-source solutions, early. Both ship from the aws/capability-insights-for-aws GitHub repo (note: the blog links the same repo URL for both Capability Insights and Workload Analysis); starting-point IAM policies live in the repo docs/ folder.
  • Private-by-design is also a friction cost. Because the dashboard has no public access, every viewer needs VPC connectivity (VPN / Direct Connect / Client VPN / SOCKS jump host) — an onboarding step the post calls out but doesn't quantify.
  • Dataset coverage depends on partition access. Beyond the public dataset, additional data sources (other partitions) are only available if the organization has access, arranged via an AWS representative.

Extracted systems / concepts / patterns

Source

Last updated · 771 distilled / 2,233 read