SYSTEM Cited by 1 source
Workload Analysis (Capability Insights)¶
Workload Analysis is the additive companion to Capability Insights for AWS (same aws/capability-insights-for-aws repo) that personalizes the Regional-availability catalog to the 20–30 services an account actually runs. It turns a Regional-expansion gap analysis from a multi-week project (evaluate 200+ services across 35+ Regions) into a ~30-minute review by filtering the catalog to a per-account "My Stuff" view. (Source: sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc)
Pipeline¶
Deployed as an additive CloudFormation stack alongside Capability Insights, it runs an AWS Step Functions state machine in three stages — a fan-out/fan-in shape (two independent analyzers in parallel, then a merge step), structurally a cousin of a scatter-gather over heterogeneous signals:
- Parallel analyzers (two branches concurrent):
- CloudTrail Analyzer — creates an AWS Glue Data Catalog table over the account's CloudTrail log bucket, then runs an Athena query (with AWS Lake Formation catalog access) extracting distinct service/API/Region/account combinations from the last N days (default 30). → what the account calls.
- CloudFormation Analyzer —
ListStacks+GetTemplateon every active stack; extractsAWS::resource types and scalar property values (strings/numbers/booleans), maps them to service names, records which stack contributed each. → what the account deploys, with per-stack attribution. - Usage Decorator — runs after both branches; reads the primary catalogs (
products.json,apis.json,cfn_resources.json) from the website bucket, intersects them with the analyzer outputs, and writes personalized files back to the bucket for the dashboard/API to consume. - Scheduled execution — an EventBridge rule triggers the state machine daily (configurable); ad-hoc runs via
POST /analysisthrough the private API Gateway or the dashboard Settings page.
The two-signal rationale¶
CloudTrail (calls) and CloudFormation (deploys) are deliberately combined: calls alone miss provisioned-but-idle resources; deploys alone miss runtime API usage. The CloudFormation side goes deeper than resource types — it surfaces which property configurations are in use and which stacks set them (e.g. AWS::EC2::Instance with t3.medium from the app stack and m5.xlarge from the data-processing stack). The merged union is a truer picture of a workload's real Regional surface.
API¶
POST /analysiswith{ "scope": "account", "analyzers": ["cloudtrail","cloudformation"], "analyzerParams": { "cloudtrail": { "bucket": "...", "daysToScan": 30 } } }.GET /capabilities?usageFilter=combined&scope=accountreturns filtered products / APIs / CloudFormation resources with usage attribution (usage.stacks,usage.properties,usage.count) andlastAnalyzedAt.- Private API Gateway — requests must originate inside the VPC.
Operational numbers¶
- First run completes in 2–5 minutes depending on account size and active-stack count.
- Worked example: a "typical web application" account filtered from 200+ services to 28, each with per-stack attribution.
- Prereq: active CloudTrail with logs in an S3 bucket; deploy with
--enable-usage-analysis --cloudtrail-bucket <bucket>.
Seen in¶
- sources/2026-10-02-aws-deploy-open-source-regional-availability-tools-in-your-vpc — the Step Functions parallel-analyzer + Usage-Decorator pipeline, the calls-vs-deploys two-signal design, and the catalog-to-28-services personalization that collapses gap-analysis effort.
Related¶
- systems/capability-insights-for-aws — the base dashboard it personalizes
- systems/aws-step-functions — the orchestration substrate
- concepts/scatter-gather-query — the fan-out/fan-in structural analogue