How AgentFlo built AI sales agents with Amazon Bedrock AgentCore - Part 2¶
Summary¶
Part 2 of AWS's two-part Architecture Blog series on AgentFlo (Salesflo's agentic-commerce platform — always-on WhatsApp/web sales agents on Amazon Bedrock AgentCore, serving merchants with $300B+ in annual transacted value). Where Part 1 covered Velocity, Standardization, and Scalability, Part 2 covers the remaining two of the five pillars — Trust (guardrails for autonomous commercial action + real-time visibility) and Reliability (a data foundation that keeps agents grounded) — plus measurable business results and a forward roadmap (voice agents, server-side tool execution). The through-line: for agents that can act on real transactions, "the model proposes; deterministic policy / structured data decides" — safety and grounding live outside the model's reasoning loop.
Trust: guardrails for autonomous commercial action¶
The core stance: in AgentFlo trust isn't only about safe responses, it's about safe action. Agents can create carts, place orders, apply discounts, and touch customer data, so policy enforcement must sit outside the model's reasoning loop — the model proposes, deterministic policy decides. Controls span the full lifecycle across three stages (three-layer-agent-guardrails, concepts/defense-in-depth):
- Pre-request (AWS Fargate layer). Detects prompt injection and handles opt-outs before requests reach the agent. WhatsApp messages are authenticated by phone number; enterprise customers (e.g. EBM) restrict access to authorized users, restaurant deployments stay open.
- During tool execution (AgentCore layer). AgentCore Gateway verifies identity and enforces order locks; Gateway policies prevent sales agents from reaching support-only tools. Cedar policies in AgentCore Policy enforce business rules like maximum discount percentages independently of model reasoning, and Policy integrates with Amazon Bedrock Guardrails so Cedar policies can invoke prompt-attack detection, content filtering, and sensitive-information blocking at the gateway boundary.
- Post-turn privacy filters. Screen outputs to block inadvertent token disclosure and unverified price claims before customers see responses.
Infrastructure security underneath: per-session microVM isolation via AgentCore Runtime (dedicated microVMs per merchant session); fine-grained Cedar Gateway tool/data policies; IAM-based agent-to-service access; VPC integration with domain-level network restrictions so agents only reach approved endpoints; and data-residency + audit-trail compliance across jurisdictions. Secrets live in AWS Secrets Manager with OIDC-authenticated CI/CD — no credentials in agent code.
Observability rounds out trust: AgentCore Observability captures structured traces per agent turn (model latency, tool invocation sequences, token usage, error rates) into CloudWatch, where AgentFlo builds dashboards (active sessions, response times, tool-call patterns), tracks P50/P95 latency and throughput per agent type, alerts on baseline deviation, and does per-merchant / per-agent / per-conversation cost attribution.
Reliability: a data foundation that keeps agents grounded¶
"Reliable agents need reliable data" — accurate action grounded in verified data, not just accurate responses. The model reasons; structured data decides. Three layers (concepts/retrieval-augmented-generation):
State management (DynamoDB)¶
A real sales journey can span 8 hours to 3 days — a two-table DynamoDB design (Session Table + Cart Table) handles stateful, autonomous, and safe requirements together (concepts/agent-memory). Per Figure 1: at the start of each turn, AgentCore Runtime loads the last 15 messages from the Session Table; the Cart Table is loaded on demand only when intent detection flags the request as cart-related — preventing the model from generating incorrect prices and quantities (intent-gated-state-loading). The design gives session continuity via conversation replay, autonomous context loading by detected intent, and data integrity via structured ground-truth storage the model can't hallucinate over.
Knowledge base grounding (Bedrock Knowledge Bases)¶
Merchants upload business-specific data (menus, clinic policies, product specs, promotion calendars) into an Amazon Bedrock Knowledge Base on S3; agents auto-retrieve and reason over it, keeping responses grounded without any prompt engineering — a direct mitigation for price/product hallucination.
Semantic product discovery (vector embeddings)¶
Every product in AgentFlo's Aurora database gets a lightweight vector embedding (the data-architecture overview also names Amazon S3 Vectors for semantic discovery), plus auto-generated searchable tags merchants can extend — so "the pink one," "the smallest one," "the chocolate with the golden wrapper" all map to the right product via natural language rather than keyword match.
Observability & billing¶
All message interactions flow through Amazon Data Firehose to S3 so merchants can track cost per conversation vs. sales revenue — ROI attribution per agent deployment and the data foundation for continuous agent improvement.
Business impact¶
Comparing agent-assisted journeys against a control group over a 90-day early deployment:
| Metric | Improvement |
|---|---|
| Net revenue uplift | +12% |
| Customer engagement | +40% |
| Conversion rate | +15% |
| Average order value | +8% |
| Customer reactivation | +20% |
At $300B annual transacted value, single-digit percentage gains translate to billions in incremental merchant revenue. Operational wins: 24/7 coverage, consistent best-practice selling, scalable per-customer personalization across thousands of concurrent sessions, and merchant onboarding in days.
What's next (roadmap)¶
- Real-time voice agents (in pilot). AgentFlo already handles WhatsApp voice notes via a two-pass transcription: a raw first pass, then a domain-aware correction pass that fuzzy-matches against the live product catalog to recover brand names and SKUs even when partially misheard, across 90+ languages (two-pass-domain-aware-transcription). Because STT, reasoning (on AgentCore Runtime), and TTS are fully independent components, any one can be swapped without disrupting the pipeline (pipeline-component-decoupling). The next step, BidiAgent (Strands SDK + AgentCore WebRTC), supports bidirectional audio streaming, natural interruptions, and concurrent tool execution (check inventory / apply a discount while still listening). The AgentCore WebRTC protocol + Amazon Kinesis Video Streams handle peer-to-peer transport for mobile/browser without relay infrastructure.
- Server-side tool execution (experimenting). Removes client-side orchestration entirely: instead of the traditional loop (model → execute → send result → repeat), the agent makes a single Amazon Bedrock Responses API call with an AgentCore Gateway ARN passed as an MCP connector; Bedrock then autonomously discovers, invokes, and processes tools through the Gateway MCP interface entirely inside AWS with no client roundtrips (server-side-tool-execution). Early result: ~30% lower latency for specialist agents with short, focused tool loops; a 3-tool sequence (check inventory → apply discount → update cart) collapses 50+ lines of orchestration into one API call, credentials stay server-side, and onboarding new agent developers is simpler.
- Integration ecosystem expansion (ongoing). Each new platform (payment, shipping, loyalty, vertical ERPs) becomes another MCP server connector on the Gateway (patterns/central-proxy-choke-point).
Key takeaways¶
- Pick your pillars first; the AWS stack follows. Clear requirements (velocity, standardization, scalability, trust) made the stack obvious: Strands for the agent layer, AgentCore Runtime for stateful sessions, Gateway for tool routing + policy, Fargate for message ingestion. (Source: sources/2026-08-21-aws-how-agentflo-built-ai-sales-agents-with-amazon-bedrock-agentcore-part-2)
- Specialized single agents beat multi-agent complexity for most interactions (single-agent-over-multi-agent); multi-agent is reserved for clean cross-domain handoffs.
- Commerce is stateful — plan for it on day one. Retrofitting state onto a stateless agent is far more painful; AgentCore stateful sessions + DynamoDB-backed context were chosen up front (concepts/agent-memory).
- Three-layer security builds trust. Pre-turn Fargate guards, per-tool Gateway/Cedar policies, and post-turn output filters together enable safe autonomous commercial action (three-layer-agent-guardrails).
- System design supports scale. Agent customization as a SaaS layer atop shared AgentCore infrastructure serves hundreds of merchants while keeping each agent's behavior personalized.
- Serverless + AgentCore = elastic commerce. Fargate ingestion + AgentCore execution scale from normal traffic to 50× flash-sale spikes without pre-provisioning.
- The feedback loop is the product. Real conversations teach agents how to close; every prior conversation on the platform benefits new merchant deployments — a compounding loop harder to copy than any single component.
Operational numbers¶
- +12% net revenue, +40% engagement, +15% conversion, +8% AOV, +20% reactivation — over a 90-day control-group comparison.
- $300B annual transacted value; 50× flash-sale traffic spikes.
- Sessions span 8 hours to 3 days; Runtime loads the last 15 messages per turn from the DynamoDB Session Table.
- Voice: 90+ supported languages; two-pass transcription (raw + domain-aware correction).
- Server-side tool execution: ~30% lower latency for short-tool-loop specialist agents; a 3-tool sequence collapses 50+ lines of orchestration into one API call.
- Sample server-side call uses
model="openai.gpt-oss-120b"against abedrock-mantleendpoint with an MCP connector pointing at the Gateway ARN (require_approval: "never"). (Model/endpoint availability varies by Region.)
Caveats¶
- Vendor/customer architecture post, not an independent retrospective. Business-impact percentages and the $300B figure are attributed to Salesflo/AgentFlo and measured over an early 90-day window; no confidence intervals or absolute baselines are given.
- Voice agents are in pilot; server-side tool execution is experimental. The ~30% latency improvement is an early measurement for a narrow class (short, focused tool loops), not a general guarantee.
- No AgentCore internals (Runtime/Gateway/Observability latency, throughput, microVM cold-start, cost) are disclosed — mechanisms are described qualitatively. Figure 1 (per-message flow) is referenced but not reproduced here.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/how-agentflo-built-ai-sales-agents-with-amazon-bedrock-agentcore-part-2/
- Raw markdown:
raw/aws/2026-08-21-how-agentflo-built-ai-sales-agents-with-amazon-bedrock-agent-db1e440a.md
Related¶
- sources/2026-08-20-aws-how-agentflo-built-ai-sales-agents-with-amazon-bedrock-agentcore — Part 1 (Velocity / Standardization / Scalability)
- systems/bedrock-agentcore
- systems/agentflo
- systems/agentcore-observability
- systems/amazon-bedrock-guardrails
- three-layer-agent-guardrails
- intent-gated-state-loading
- server-side-tool-execution
- two-pass-domain-aware-transcription
- concepts/defense-in-depth
- concepts/agent-memory
- pipeline-component-decoupling