Skip to content

SYSTEM Cited by 2 sources

Amazon Bedrock AgentCore Evaluations

What it is

Amazon Bedrock AgentCore Evaluations is the evaluation surface of Amazon Bedrock AgentCore that scores each agent decision using built-in and custom evaluators, applying an LLM-as-a-judge approach. It is the automated quality gate that decides which agent outputs are trusted to proceed and which are escalated for human review.

Evaluation dimensions

In the clinical-trial screening architecture, AgentCore Evaluations scores every screening decision across three dimensions (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore):

  • Clinical accuracy / reasoning — built-in evaluators: correctness (against labs, diagnoses, meds), faithfulness (reasoning grounded in patient data, not invention), coherence (no logical contradictions), context relevance (right protocol and records retrieved), goal success rate (full workflow ran end to end). Custom evaluators add eligibility accuracy (each inclusion/exclusion criterion evaluated correctly) and criteria coverage (no criteria skipped, especially safety-critical thresholds).
  • Operational effectiveness — helpfulness, conciseness, relevance, instruction following (expected structured format), and tool selection / parameter accuracy.
  • Safety compliance — strictest thresholds. Harmfulness detection, stereotyping detection, and custom evaluators for safety-flag detection (every significant concern surfaced — a single miss is treated as critical) and uncertainty acknowledgment (recommend human review on ambiguous data rather than an overconfident call).

How it gates the pipeline

Decisions that pass with high confidence proceed to the clinician dashboard. Decisions that fall below quality thresholds are routed to human review with the specific evaluation concern highlighted — implementing llm-judge-gated-human-review-routing. Continuous sampling detects drift, and CloudWatch alerts teams when quality drops below thresholds, providing ongoing evidence the agent performs within validated parameters (supporting GxP with minimal manual testing). Clinician overrides feed back into the evaluators' ground-truth dataset (continuous-evaluation-feedback-loop). (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)

Seen in

  • sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore — LLM-as-judge scoring layer that gates which screening decisions reach clinicians
  • sources/2026-08-26-aws-closing-the-ai-agent-trust-gap-with-graduated-autonomy — the evaluation engine behind the delivery gate in graduated autonomy: run in an AWS CodePipeline stage against ground-truth fixtures (including adversarial prompt-injection / data-exfiltration cases) in staging; a single unauthorized tool call fails the release (delivery-gate-blocks-degraded-agent-version).
  • sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore — a two-layer evaluation design over an inventory pipeline. Three online evaluators run every session at 100% sampling during rollout: Builtin.GoalSuccessRate (session), Builtin.Helpfulness (trace), and a custom 3-point ConstraintCompliance LLM-as-a-Judge rubric (1.0 silent violation / 2.0 flagged / 3.0 compliant) that deliberately rewards flagging a violation over hiding it. A fourth — a custom code-based Lambda ForecastAccuracyEvaluator (WAPE, signed bias, pinball loss @P90, P10–P90 coverage) — runs on-demand, not online, because ground truth arrives one horizon later (registering it online would produce HORIZON_MISMATCH on every call). Illustrates the evaluator-type choice as a mirror of patterns/deterministic-tool-vs-llm-judgment: forecast accuracy is arithmetic (code-based); agent behavior is judgment (LLM-as-a-judge). Evaluators are registered via the AgentCore control plane and referenced by ARN; online configs use create_online_evaluation_config with a samplingConfig.
Last updated · 766 distilled / 2,225 read