Closing the AI agent trust gap with graduated autonomy¶
Summary¶
An AWS Architecture Blog piece that introduces graduated autonomy, an architectural pattern for governing what an AI agent is allowed to do as a function of its demonstrated reliability. The central observation is that classic IAM answers "who can do what?" once, at provisioning, assuming the principal behaves consistently — an assumption a large language model breaks, because the same agent can be accurate Monday and hallucinating Tuesday after a prompt change or model update. The distance between what an agent could do and what an operator trusts it to do is the agent's trust gap. Graduated autonomy closes the gap with a six-layer closed loop: a scoring engine computes a rolling trust score, a tier system translates sustained scores into autonomy levels, a pre-execution layer blocks dangerous actions in-process, an enforcement layer applies tiers via Cedar policies outside the agent, a post-execution layer captures audit records and feeds outcomes back to scoring, and a delivery gate keeps degraded agent versions out of production. Every layer is configuration, not code, and each embodies one deliberate design decision. Built on Amazon Bedrock AgentCore (runtime, Gateway, Policy, Evaluations), DynamoDB for trust state, and AWS CodePipeline for the delivery gate.
Key takeaways¶
-
The trust gap is a consistency problem, not an access problem. IAM grants are static — "the same agent can be accurate Monday and hallucinated Tuesday after a prompt change or model update." Most teams collapse the decision into a binary: full access (risky — agents fail unpredictably) or read-only (safe but leaves most of the agent's value unused). Graduated autonomy replaces the binary with a continuum earned through sustained reliability (Source: sources/2026-08-26-aws-closing-the-ai-agent-trust-gap-with-graduated-autonomy). See agent-trust-gap.
-
Closing the gap needs three capabilities raw API logs don't give: visibility, decision provenance, and reversibility. Visibility = whether an action was safe, not just what happened. Decision provenance = tracing an action back to the signal that triggered it, the alternatives considered, and the confidence held (decision-provenance). Reversibility = pre-action state capture so operators can recover from incorrect actions (pre-action-state-capture).
-
Scoring engine: five weighted dimensions over a rolling 50-action window. Accuracy 25%, Safety 20%, Consistency 20%, Compliance 20%, Efficiency 15% — composite score 0–100. Crucially, safety is an independent floor, never averaged away by strong metrics: a dangerous individual dimension can't be hidden by strength elsewhere (safety-as-independent-floor, rolling-window-trust-score).
-
Tier system: start every agent at T1; promote slowly, demote immediately. T1 Probation (0–40): read/list only, two tools visible. T2 Supervised (41–70): add writes, human approves high-risk. T3 Trusted (71–90): execute/modify, anomalies flagged. T4 Autonomous (91–100): full access, post-hoc audit only. Every new agent starts at T1 regardless of test performance (patterns/staged-rollout).
-
Hysteresis prevents tier flapping. Promotion into a tier requires a score 5 points above that tier's range floor; demotion happens at the floor. An agent sitting at a boundary cannot oscillate (tier-hysteresis). Promotion also requires the score to stay above threshold for the entire rolling window; demotion is immediate when safety drops below its floor or injection is detected.
-
Pre-execution layer: fast in-process filters, backstopped, never solely trusted. Six independently-scored signals evaluate every tool call before it runs — adversarial-injection detection (one match = instant block + trust penalty), sensitive-target regex (credentials/tokens/keys in args), dangerous-tool detection, behavioral consistency vs. historical tool-use distribution, confidence calibration (overconfident failures penalized at 2×), and reasoning-quality (did the agent reason before acting). These are first-pass filters; the enforcement layer's deny-by-default policies catch anything they miss.
-
Enforcement layer: deny by default, enforced outside the agent's process. AgentCore Gateway routes every MCP tool invocation through Policy in AgentCore, which evaluates Cedar policies with forbid-wins semantics — one satisfied forbid overrides any number of permits. Tier maps to policy state: Probation applies a forbid on write/execute/delete; promotion removes the forbid; demotion re-applies it. In enforce mode the Gateway lists only tools policy could permit, so blocked tools never appear — "the agent is unlikely to call a tool it has never seen." Listing is a meta-action; each invocation is still evaluated with full request context. Bedrock Guardrails complements Policy for model-level content safety (out-of-band-agent-enforcement).
-
Post-execution layer: the Think → Plan → Act → Observe → Score audit record makes recovery possible. Every action logs its reasoning chain (Think), tool + input + pre-execution score (Plan), Cedar match + Gateway route (Act), success/failure + output data (Observe), and trust impact + per-dimension scores + tier change (Score). Plan and Act capture pre-action state — the substrate for reversibility. A provenance query endpoint answers plain-English questions about any decision; entries persist to DynamoDB (think-plan-act-observe-score-audit-record).
-
Delivery gate: one unauthorized tool call in adversarial tests blocks release. Each change to the agent's prompt, config, or tool definitions triggers an AWS CodePipeline run that deploys the candidate to staging and runs it against ground-truth fixtures with AgentCore Evaluations. Fixtures include adversarial cases (prompt injection, data exfiltration); a single unauthorized tool call in any adversarial case fails the gate. The version that passes becomes the last known stable version (delivery-gate-blocks-degraded-agent-version).
-
Production monitoring, recovery, and multi-agent safety. Synthetic honeypot cases with known expected behavior are injected into a small share of traffic; validation checks the tool-call trajectory (expected tools, order, no forbidden tools) rather than nondeterministic NL output, and honeypot results stay out of production metrics (honeypot-trajectory-validation). When safety drops below the floor, demotion narrows permissions and the framework redeploys the last known stable version — restoring known-good code alongside a tighter permission set. An emergency stop pushes a single Cedar deny-all policy that (typically within seconds) denies all tool invocations with no redeployment (emergency-stop-deny-all). Operator judgment feeds scoring directly: the rolling rate at which operators reject proposed actions caps the effective safety metric (30% rejections cap safety at 70). In multi-agent systems, a delegated action's effective tier is the minimum across the delegation chain, closing the delegation privilege-escalation path (delegated-tier-minimum).
Operational numbers¶
- Rolling window: 50 actions per agent for the trust score.
- Scoring weights: Accuracy 25% · Safety 20% · Consistency 20% · Compliance 20% · Efficiency 15% (0–100 composite).
- Tier bands: T1 0–40 · T2 41–70 · T3 71–90 · T4 91–100.
- Hysteresis: +5 points above range floor to promote; demote at floor.
- Confidence penalty: overconfident failures penalized at 2× normal.
- Operator-rejection cap: 30% rejection rate → effective safety capped at 70.
- Delivery gate: 1 unauthorized tool call in any adversarial fixture fails the release.
- DynamoDB lookup: current-tier read on every invocation, typically single-digit milliseconds.
- Emergency stop: Cedar deny-all active typically within seconds, no redeploy.
Caveats¶
- Reference architecture, not a benchmarked production system. The weights, tier bands, and thresholds are offered as a starting template ("take the dimension weights and tier boundaries … as a starting template for one agent in your fleet"), not values validated against a specific workload. No throughput, false-positive, or latency-of-scoring numbers are given.
- Pre-execution injection detection is pattern-matching. The post is explicit that these are fast first-pass filters, not a complete defense — the deny-by-default enforcement layer is the real backstop.
- Trust score is a lagging signal. A 50-action rolling window means a newly-degraded agent keeps its tier until enough bad actions accumulate — mitigated by immediate demotion on a safety-floor breach or injection hit, and by the delivery gate catching regressions before they reach production.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/closing-the-ai-agent-trust-gap-with-graduated-autonomy/
- Raw markdown:
raw/aws/2026-08-26-closing-the-ai-agent-trust-gap-with-graduated-autonomy-91e6670d.md
Related¶
- graduated-autonomy
- agent-trust-gap
- rolling-window-trust-score
- tier-hysteresis
- safety-as-independent-floor
- systems/agentcore-policy
- systems/bedrock-agentcore
- systems/aws-codepipeline
- out-of-band-agent-enforcement
- concepts/audit-trail