Skip to content

SYSTEM Cited by 1 source

Workload Analysis (Capability Insights)

Workload Analysis is the additive companion to Capability Insights for AWS (same aws/capability-insights-for-aws repo) that personalizes the Regional-availability catalog to the 20–30 services an account actually runs. It turns a Regional-expansion gap analysis from a multi-week project (evaluate 200+ services across 35+ Regions) into a ~30-minute review by filtering the catalog to a per-account "My Stuff" view. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)

Pipeline

Deployed as an additive CloudFormation stack alongside Capability Insights, it runs an AWS Step Functions state machine in three stages — a fan-out/fan-in shape (two independent analyzers in parallel, then a merge step), structurally a cousin of a scatter-gather over heterogeneous signals:

  1. Parallel analyzers (two branches concurrent):
  2. CloudTrail Analyzer — creates an AWS Glue Data Catalog table over the account's CloudTrail log bucket, then runs an Athena query (with AWS Lake Formation catalog access) extracting distinct service/API/Region/account combinations from the last N days (default 30). → what the account calls.
  3. CloudFormation Analyzer — ListStacks + GetTemplate on every active stack; extracts AWS:: resource types and scalar property values (strings/numbers/booleans), maps them to service names, records which stack contributed each. → what the account deploys, with per-stack attribution.
  4. Usage Decorator — runs after both branches; reads the primary catalogs (products.json, apis.json, cfn_resources.json) from the website bucket, intersects them with the analyzer outputs, and writes personalized files back to the bucket for the dashboard/API to consume.
  5. Scheduled execution — an EventBridge rule triggers the state machine daily (configurable); ad-hoc runs via POST /analysis through the private API Gateway or the dashboard Settings page.

The two-signal rationale

CloudTrail (calls) and CloudFormation (deploys) are deliberately combined: calls alone miss provisioned-but-idle resources; deploys alone miss runtime API usage. The CloudFormation side goes deeper than resource types — it surfaces which property configurations are in use and which stacks set them (e.g. AWS::EC2::Instance with t3.medium from the app stack and m5.xlarge from the data-processing stack). The merged union is a truer picture of a workload's real Regional surface.

API

  • POST /analysis with { "scope": "account", "analyzers": ["cloudtrail","cloudformation"], "analyzerParams": { "cloudtrail": { "bucket": "...", "daysToScan": 30 } } }.
  • GET /capabilities?usageFilter=combined&scope=account returns filtered products / APIs / CloudFormation resources with usage attribution (usage.stacks, usage.properties, usage.count) and lastAnalyzedAt.
  • Private API Gateway — requests must originate inside the VPC.

Operational numbers

  • First run completes in 2–5 minutes depending on account size and active-stack count.
  • Worked example: a "typical web application" account filtered from 200+ services to 28, each with per-stack attribution.
  • Prereq: active CloudTrail with logs in an S3 bucket; deploy with --enable-usage-analysis --cloudtrail-bucket <bucket>.

Seen in

Last updated · 771 distilled / 2,233 read