Skip to content

From zero-shot forecast to purchase order with Amazon Bedrock AgentCore

Summary

An AWS Architecture Blog (Learning level 300, Best Practices) reference architecture that turns raw demand forecasts into audited purchase orders by composing two complementary abstractions: Amazon Chronos2, a time-series foundation model doing zero-shot probabilistic forecasting on SageMaker Serverless Inference, and a four-agent orchestration (Supervisor, Preprocessing, Forecasting, Reporting) built on the Strands Agents SDK and deployed on Amazon Bedrock AgentCore. The organizing thesis is a crisp design test — "if I fix the input, will the output always be the same?": if yes, implement it as a deterministic @tool; if no, the agent's LLM reasoning is the logic. A second boundary test ("if the agent calls this tool wrongly, does the mistake propagate externally or stop at the agent?") decides which of eight tools sits behind the Gateway with Cedar policy — exactly one (save_decision). The post also maps each of AgentCore's six sub-services to a single production concern and layers two evaluation cadences (online behavior scoring + on-demand forecast-accuracy scoring).

Key takeaways

  1. Zero-shot forecasting removes the per-SKU ML pipeline. Classical methods (ARIMA, Holt-Winters) and gradient-boosting/deep approaches (LightGBM, DeepAR, TFT) require per-SKU model fitting; a 10,000-SKU catalog means 10,000 models to train, validate, and retrain. Chronos2 (an encoder-only T5-style transformer pre-trained on diverse time series) forecasts a new SKU with no training job — historical window in, probabilistic forecast out — cutting per-SKU onboarding from 2–3 weeks of model training to under 5 minutes of CSV upload (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  2. The agent-vs-tool split is the most consequential design decision. The practical test: "If I fix the input, will the output always be the same?" Yes → @tool (LLM calls it, function does the work). No → the LLM reasoning is the logic. Applying it to 13 components, only 6 are agent-reasoning (parse request, select covariates, interpret anomalies, generate rationale, retry-or-escalate decision, summarize); the rest are deterministic tools. Bedrock inference happens in exactly four places — the four agents. This keeps per-run LLM cost bounded and reasoning quality high (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  3. A second boundary — in-process vs Gateway — decides trust, not determinism. Test: "If the agent hallucinates and calls this tool wrongly, does the mistake propagate to external systems or stop at the agent's memory?" Of eight tools, exactly one (save_decision, the authoritative order write) sits behind the Gateway with Cedar authorization. generate_forecast_chart also writes S3 but stays in-process because a chart is a visualization, not a decision of record. "That proportion is the norm, not the exception" (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  4. Multi-agent beats monolithic for context isolation and cost. The implementation uses the Agents-as-Tools pattern: the Supervisor is a single Strands agent whose tool list is the three specialists, each wrapped as a @tool. Each specialist @tool invocation spawns a fresh Strands agent with its own context window, system prompt, and tool subset; results return as a compressed CLUES_FORMAT labeled-output envelope, so the Supervisor sees labeled deltas rather than full reasoning transcripts. max_node_executions=10 bounds the loop (base sequence + 3 retries + safety stop) (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  5. Constraint violation is a first-class workflow state, not an error. When validate_constraints returns approved: false, its violations array carries a human-readable string per violated constraint. The Supervisor interprets it and either re-invokes Forecasting with adjusted constraints or escalates to the user — tracking iteration count in AgentCore short-term memory and stopping after 3 iterations to avoid silent infinite loops. In the walkthrough, a $500 budget cap on a $9,412.50 order causes the agent to surface three options (raise budget / accept shortfall / delay promotion) rather than silently truncating — human-in-the-loop (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  6. Scale-to-zero drives a 98% inference-cost reduction. SageMaker Serverless Inference for the Chronos2 endpoint costs ~$15/month (≈500 invocations × 8s each) vs ~$1,091/month for an always-on ml.g5.2xlarge GPU endpoint — a 98% reduction — at the cost of a 30–60s cold start, which is acceptable for nightly/weekly batch forecasting. The AgentCore Runtime follows the same model: per-session microVM, up to 8-hour session duration, no idle cost. Break-even toward always-on is ~a few thousand invocations/month (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  7. Structured JSON contracts between agents prevent the top multi-agent failure mode. Each agent outputs a typed JSON structure the next consumes — explicit about which covariates were used (auditable Preprocessing decision), carrying order_decision, anomaly_flags, and a first-class rationale field. "Implicit coupling through unstructured text — where one agent returns a paragraph and the next tries to extract numbers from it — is the most common failure mode in multi-agent systems" (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  8. Two evaluation layers: behavior (online) + accuracy (on-demand). Three online evaluators (Builtin.GoalSuccessRate, Builtin.Helpfulness, and a custom 3-point ConstraintCompliance LLM-as-a-Judge rubric: 1.0 silent violation / 2.0 flagged / 3.0 compliant) run every session at 100% sampling during rollout. A fourth — a code-based Lambda ForecastAccuracyEvaluator (WAPE, signed bias, pinball loss @P90, P10–P90 coverage) — runs on-demand because ground truth arrives one horizon later. The evaluator-type choice follows the same logic as agent-vs-tool: forecast accuracy is arithmetic (code-based), agent behavior is judgment (LLM-as-a-judge) (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

  9. One boundary per AgentCore service. Runtime = where agents run (microVM, scale-to-zero); Gateway + Policy = what writes reach external systems (Cognito JWT + Cedar); Memory = what the agent carries across sessions (semantic + user-preference strategies, no summary); Observability = why a decision was made (auto-traced to CloudWatch); Evaluations = whether the agent still behaves after deployment. "AgentCore is not a monolithic platform you adopt or refuse — it is a set of services, each addressing one specific concern" (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

Operational numbers

  • Forecast accuracy: median WAPE 12.3% (P50 vs actuals) across 50 SKUs over a 4-week horizon (internal testing).
  • Onboarding: per-SKU time cut from 2–3 weeks (model training) to under 5 minutes (CSV upload).
  • Cost: monthly inference ~$1,091 (always-on GPU) → ~$15 (Serverless) — 98% reduction.
  • Latency: ~8 s end-to-end per SKU (excluding cold start); 30–60 s Serverless cold start.
  • Throughput: 10,000-SKU nightly run in under 30 minutes at ~100-way parallelism, bounded by SageMaker Serverless concurrency quota.
  • Retry loop: capped at 3 iterations, then escalate; max_node_executions=10.
  • Cedar high-value deny threshold: context.input.budget_used > 50000 is denied regardless of principal.
  • Production targets: P95 end-to-end latency < 90 s (batch session); ConstraintCompliance ≥ 2.0 on 95% of sessions; rolling 4-week WAPE < 20% on top-20 SKUs by revenue. Error budget = 5% of sessions below 2.0.
  • Accuracy thresholds: WAPE < 15% (M5 "strong baseline"), |bias| < 0.05, coverage 0.75–0.85 → ACCURATE; pinball @P90 penalizes under-coverage 9× (asymmetric stock-out cost).
  • Models: Claude Sonnet 4.5 (us.anthropic.claude-sonnet-4-5-20250929-v1:0) for all four agents; Chronos2 for forecasting.

Systems / concepts / patterns extracted

Caveats

  • WAPE 12.3% is from internal testing across 50 SKUs over 4 weeks, not a production fleet at 10,000-SKU scale. Thresholds (WAPE, coverage, pinball) are described as retail-demand defaults to tune per catalog.
  • Cost figures are us-east-1 estimates assuming ~500 inference calls/month; actual costs vary by Region and usage pattern.
  • The architecture is a reference/best-practices post with code snippets, not a disclosed production system with a post-incident track record.
  • Serverless Inference constraints elsewhere (no GPU, 6 GB memory cap — see systems/aws-sagemaker-endpoint) are not emphasized here because the Chronos2 payload sits well under a megabyte and the batch duty cycle tolerates cold starts.

Source

Last updated · 766 distilled / 2,225 read