Skip to content

AWS 2026-08-19

Read original ↗

AI-powered clinical trial eligibility and safety using Amazon Bedrock AgentCore

Summary

AWS presents a proposed architecture for a Clinical Trial Eligibility and Safety Agent that automates patient screening while keeping clinicians as the decision-makers through human-in-the-loop (HITL) oversight. The system ingests fragmented patient records (FHIR notes, labs, imaging, medication histories) through AWS HealthLake, assembles them into a knowledge graph of patients/molecules/endpoints/markets, and runs three specialized agents on Amazon Bedrock AgentCore to pre-screen, evaluate against protocol inclusion/exclusion criteria, and match to sites. Every decision is scored by an LLM-as-a-judge evaluation layer (AgentCore Evaluations); low-confidence or safety-flagged cases route to a tiered clinician review queue. Clinician overrides are captured as corrected ground truth that feeds back into evaluation and prompt/retrieval tuning. Immutable audit records land in DynamoDB for FDA 21 CFR Part 11 traceability, and CloudWatch provides end-to-end observability. The post targets solution architects and life-sciences engineering teams; it is a reference architecture, not a production retrospective. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)

Key takeaways

  1. The problem is evidence fragmentation, not data scarcity. Eligibility and safety signals are scattered across EHR notes, lab portals, imaging reports, and medication histories, forcing study teams to manually reconstruct each candidate's clinical picture. The agent organizes evidence; it does not supply missing data. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  2. A knowledge graph replaces repeated cross-source joins. Patients, molecules, endpoints, and markets are nodes; relationships are edges. To answer an eligibility/safety question the agent traverses edges (diagnosis → associated labs; medication → known interactions) rather than re-querying and joining disconnected sources each time. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  3. Three agents, each scoped to one pipeline phase. A pre-screening agent resolves threshold gates (consent validity, high-level profile fit, washout completion); a detailed screening agent walks every inclusion/exclusion criterion by retrieving the right FHIR resource (Observation for labs, Condition for diagnoses, MedicationStatement for meds) and emits a structured determination (Eligible / Ineligible / Requires Review) with a per-criterion evidence matrix, confidence scores, and a reasoning summary citing source records; a site & enrollment agent matches cleared patients to sites by proximity, capability, and open capacity. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  4. AgentCore Runtime supplies the shared substrate. Agents run in the AgentCore Runtime, connect to tools through MCP Gateway (systems/agentcore-gateway), maintain session memory (systems/agentcore-memory) so they reference earlier findings without re-querying, enforce identity-based least-privilege data access (systems/agentcore-identity), and use a built-in code interpreter for dynamic clinical calculations (eGFR, BMI). (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  5. Guardrails wrap all three agents. Amazon Bedrock Guardrails enforce PII/PHI filtering, content-safety controls, grounding checks (anchor responses in retrieved evidence, not model parametric knowledge), and denied-topic boundaries to keep agents within screening scope. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  6. An LLM-as-judge layer decides which cases clinicians see. AgentCore Evaluations scores every decision across three dimensions — clinical accuracy (correctness, faithfulness, coherence, context relevance, goal success), operational effectiveness (completeness, conciseness, tool selection/parameter accuracy, instruction following), and safety compliance (custom evaluators verify no safety-critical criterion was skipped, uncertainties are acknowledged, and flags route to the right tier). High-confidence passes proceed to the dashboard; sub-threshold cases route to human review with the specific concern highlighted. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  7. Human review is tiered by complexity, and the clinician is always the decision-maker. Flagged cases flow into a PI review queue, a study-coordinator dashboard, patient communication, and escalation to a medical director for high-risk edge cases. Clinicians retain complete override capability at every stage; routine high-confidence checks proceed automatically, and cases unreviewed beyond set timeframes escalate automatically. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  8. Overrides become ground truth — a continuous learning loop. When a clinician overrides a recommendation, the system captures the corrected decision plus the clinician's reasoning; these expand the ground-truth dataset used by AgentCore Evaluations and surface patterns that inform prompt and retrieval tuning. (continuous-evaluation-feedback-loop) (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  9. Auditability and compliance are first-class, layered concerns. Immutable audit records in DynamoDB capture clinician ID, timestamp, patient/trial IDs, outcomes, AI recommendations, and full workflow execution history, designed to support FDA 21 CFR Part 11. Security is layered: HealthLake is HIPAA-eligible with encryption and SMART on FHIR authorization; Bedrock is HIPAA-eligible and never shares data with model providers; PrivateLink keeps traffic off the public internet; AgentCore enforces declarative authorization policies and runs inside a VPC; CloudTrail records API calls. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)
  10. The framework generalizes beyond screening. The same agent orchestration, evaluation pipeline, and compliance infrastructure are positioned to support future post-enrollment agents (adverse-event detection, protocol-deviation tracking, retention-risk prediction, re-screening triggers), each inheriting the existing scoring, logging, and auditability without a separate governance framework. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)

Architecture

Step 1  Ingestion   AWS HealthLake normalizes EHR / labs / imaging / meds → FHIR R4
                          │
                          ▼  evidence assembled into knowledge graph
Step 2  Orchestration   Amazon Bedrock AgentCore Runtime
        ┌───────────────────────────────────────────────────────────┐
        │ Pre-screening agent   → consent / profile / washout gates   │
        │ Detailed screening    → per-criterion FHIR retrieval + eval │
        │ Site & enrollment     → site match by proximity/capacity    │
        └───────────────────────────────────────────────────────────┘
          tools via MCP Gateway · session memory · identity least-privilege
          all agents behind Amazon Bedrock Guardrails (PII/PHI, grounding,
          content safety, denied topics); Knowledge Bases hold protocols
                          │
                          ▼
Step 3  Evaluation   AgentCore Evaluations (LLM-as-judge)
                     clinical accuracy · operational effectiveness · safety
                     pass (high confidence) ─────────────┐
                     flag (sub-threshold / safety) ──┐    │
                          │                           │    │
Step 4  HITL review   PI queue → coordinator dashboard │    │  clinician
                      → medical-director escalation ◄──┘    │  dashboard
                      clinician override captured ──────────┘
                          │  corrected decision + reasoning
                          ▼  feeds ground truth → prompt/retrieval tuning
Step 5  Observability & audit
        CloudWatch: agent traces, latency, error rates, judge scores,
        HITL override/review-latency metrics, alarm-based escalation
        DynamoDB immutable audit trail (FDA 21 CFR Part 11)
        CloudTrail API-call log · PrivateLink · VPC isolation

The credibility of the pipeline rests on two layers: the automated evaluation layer that scores every decision, and the HITL layer that gives clinicians final authority. LLM-as-Judge decides which cases clinicians see and how they're prioritized; the HITL workflow decides how clinicians act. Together they form a loop where human judgment both safeguards and improves agent performance. (Source: sources/2026-08-19-aws-ai-powered-clinical-trial-eligibility-and-safety-using-amazon-bedrock-agentcore)

Operational evidence and constraints

Item Article disclosure Design consequence
Motivating problem 80% of clinical trials miss enrollment timelines; each day of delay ≈ $500,000 (cited external studies) Latency of manual chart review is the bottleneck the architecture targets; claim is "days to minutes" for patient matching
Eligibility scarcity Fewer than 5% of cancer patients enroll under strict oncology criteria (cited) Consistent, multi-step criteria interpretation across sites is a stated goal
Confidence threshold Set by the clinician at trial onset HITL routing is a tunable business control, not a fixed model property
Safety-flag miss Treated as critical — a single missed safety-critical criterion is unacceptable Custom safety evaluators + guardrail grounding checks are load-bearing, not optional
Compliance posture HIPAA, FDA 21 CFR Part 11, GxP, GDPR named; "consult your compliance team and conduct your own assessment" The post provides architecture, not certification; audit trail is designed to support Part 11, not guarantee it
Maturity Proposed reference architecture; no production metrics, latency numbers, or accuracy rates for the system itself are disclosed Numbers above are external motivating figures, not measured results of this system

Caveats

  • Reference architecture, not a production case study. All quantitative figures (enrollment-miss rate, cost-per-day, <5% oncology enrollment) are external citations motivating the design; the post discloses no measured latency, accuracy, or throughput for the system itself.
  • Marketing tail. The post closes with a "schedule a 30-minute architecture review" call to action. The bulk (>60%) of the body is genuine serving-infra architecture (ingestion, multi-agent orchestration, evaluation, HITL, audit, observability, security layering), which is why it passes scope.
  • Clinical/regulatory claims are the customer's responsibility. The audit trail is designed to support FDA 21 CFR Part 11; the post explicitly defers compliance validation to the reader.

Source

Last updated · 766 distilled / 2,225 read