Skip to content

AWS 2026-09-30 Tier 1

Read original ↗

How MHK built a HIPAA-eligible agentic AI solution on Amazon Bedrock

Summary

MHK (a Hearst Health company, #1 in payer care-management in the 2024 Best in KLAS report) built the SmartProminence AI Orchestrator — a multi-tenant, HIPAA-eligible agentic workflow framework running entirely on AWS — to replace the unsustainable practice of building a separate AI system (with its own compliance infra, audit trails, and security controls) for each medical-management workflow (medical, pharmacy, grievance, appeals). The design centers on a controller-agent pattern: a stateless Workflow Engine Controller resolves a version-pinned workflow definition into a DAG (via Kahn's algorithm), dispatches individual steps to LLM processing agents depth-by-depth, and all components communicate asynchronously through Amazon SQS queues. An Agent Orchestration Core (Spring Boot on Fargate) is the only component that touches the database; controllers and agents reach it through its REST API, which also proxies all LLM calls to Amazon Bedrock (Claude). The result: 90% reduction in manual review effort and new AI features shipping in ~2 weeks instead of 3–6 months, with a single orchestrator-level HIPAA/SOC 2 attestation that every new agent inherits automatically.

Key takeaways

  1. Compliance as an orchestrator-level concern, not a per-feature concern. The central insight: instead of attesting each AI feature independently, MHK maintains one HIPAA/SOC 2 attestation at the orchestrator level; new agents inherit per-client encryption, audit logging, token-based access, and data minimization automatically. This is what collapsed the 3–6-month release cycle to ~2 weeks. (Source: article "Results and impact" + "Conclusion".)

  2. Controller-agent split = orchestrator/worker for LLM workflows. The Workflow Engine Controller decides what to do and in what order; LLM processing agents execute individual steps. This separation makes the workflow the single source of truth, lets agent processing scale independently from workflow logic, and keeps agents operating only on specific, actionable steps. (Source: "Controller-agent pattern with Amazon Bedrock".)

  3. DAG-based depth-parallel execution with conditional skips. On job arrival the controller loads a version-pinned workflow definition, resolves step dependencies into a DAG with Kahn's algorithm, pre-creates step executions in a WAITING state, and dispatches agents layer by layer — steps at the same depth run in parallel; the controller polls for completion before advancing. A Spring Expression Language (SpEL) conditional evaluated against completed steps' structured outputs can skip downstream steps (e.g. skip pharmacy verification if the case has no prescription drugs), saving time and token cost. (Source: "DAG-based parallel execution".)

  4. Fully event-driven; the only blocking call is the Bedrock LLM invocation. The core enqueues a ControllerTaskMessage on the Controller Invoke Queue to start a workflow and an AgentTaskMessage on the Agent Invoke Queue to dispatch a step. SQS gives four properties: stateless processing (any instance picks up pending messages), independent horizontal scaling of agents via ECS Fargate, fault isolation (a failed task doesn't block parallel steps; dead-letter queues capture failures), and decoupled deployment (ship new agent versions with no system-wide restart). "The system can process hundreds of concurrent jobs without resource contention." (Source: "Event-driven orchestration with Amazon SQS".)

  5. Java virtual threads for intra-step parallelism. Within a single step, agents spawn a Java 21 virtual thread per execution to process multiple items concurrently — e.g. extracting data from every page of a multi-page document simultaneously. (Source: "Controller-agent pattern" — agent pipeline.)

  6. Dynamic agent registry = configuration-driven, not engineering-driven. To add an agent type, a developer defines prompt templates, input/output schemas, and model selection, then registers it through the core's API; Terraform auto-provisions the supporting SQS queues, IAM roles, and ECS task definitions. The agent is immediately available for workflow steps with no new compliance certification because it runs inside the already-certified orchestrator. Engineers focus on prompt design and workflow logic rather than infra scaffolding. (Source: "Dynamic agent registry".)

  7. Capability-token model enforces per-dispatch least privilege. A controller-scoped token can read workflow definitions/job data, dispatch agent tasks, and create step executions; an agent-scoped token can only read its step's input, write its own result, call the LLM through the proxy, and upload artifacts. Tokens are generated per-dispatch and signed/validated through AWS KMS, so even a compromised agent cannot reach data from other steps, workflows, or clients. (Source: "Capability token model".)

  8. Multi-tenant data isolation via per-client KMS keys + layered encryption. Every client has its own KMS key. S3 documents are double-encrypted (S3 SSE + client-specific KMS); the database adds row-level encryption on top of RDS storage-level encryption. A misrouted job cannot be decrypted by the receiving agent because it lacks that client's KMS key — tenant isolation enforced cryptographically, not just by routing. (Source: "Per-client encryption".)

  9. Application-layer responsible-AI validation beyond general guardrails. Each agent post-processes model responses against expected output schemas, cross-references extracted data against source documents to detect hallucinations, rejects responses below a confidence threshold, and verifies outputs reference only the patient's own clinical records / policy-specific criteria. LLM request/response bodies are never logged — only token counts and content hashes — satisfying compliance and reproducibility. (Source: "Responsible AI controls" + "Compliance controls".)

  10. Conversational memory via durable structured outputs, not re-ingestion. Each execution returns a job ID the upstream system associates with a case; over a case's life (intake, policy review, appeal — 3+ executions) each produces structured outputs that stay available in S3 (encrypted with the client's KMS key) for later executions. Once a 30-page record is analyzed, its structured output is reused without re-running ingestion. (Source: "Conversational memory and case association".)

Architecture (as described)

  • Agent Orchestration Core — Spring Boot on AWS Fargate; exposes a REST API for job submission, workflow management, LLM proxying, and token management. Only component with direct database access → strict data-access boundary. All Bedrock calls go through its proxy.
  • Workflow Engine Controller — resolves the version-pinned workflow DAG (Kahn's algorithm), pre-creates WAITING step executions, evaluates SpEL conditions, and dispatches agents depth-by-depth, polling for completion between depths.
  • LLM processing agents — stateless; standardized pipeline: input binding (resolve expressions over prior step results) → optional vision processing of scanned docs → prompt assembly with enriched context → Bedrock (Claude) invocation → post-processing (field extraction, type coercion, structured output, hallucination/confidence checks). Use Java 21 virtual threads for per-item parallelism within a step.
  • Messaging — SQS Controller Invoke Queue (ControllerTaskMessage) + Agent Invoke Queue (AgentTaskMessage); each message carries a capability token scoped to that operation's data; dead-letter queues capture failures.

AWS services used (from the article's table)

Service Role
systems/amazon-bedrock Foundation-model inference (Claude), IAM role-based auth
systems/amazon-ecs Containerized orchestration core, controllers, agents (on Fargate)
systems/aws-sqs Event-driven inter-component messaging + DLQs
systems/aws-rds Workflow definitions, execution state; MySQL 8.4, Multi-AZ
systems/aws-s3 Job artifacts, document storage, immutable configs (KMS-encrypted)
systems/aws-kms Capability-token signing + per-client encryption keys
systems/amazon-cognito OAuth2/JWT API authentication + user identity
systems/aws-cloudwatch Logging, metrics, token-usage tracking, alerting (no PHI)
Elastic Load Balancing TLS 1.3-terminated application load balancer
Amazon VPC Network isolation; private subnets; VPC endpoints for AWS service access

Operational numbers

  • 90% reduction in manual review effort: document intake dropped from 5–10 min (manual) to under 1 minute (automated + human-in-the-loop verification); complex medical-director reviews reduced from hours of document searching to a ~30-second approve/deny decision.
  • 85% faster AI feature deployment: 3–6 month release cycle → ~2 weeks.
  • Unified compliance: single orchestrator-level HIPAA + SOC 2 attestation covering all agents.
  • RDS runs MySQL 8.4, Multi-AZ; TLS 1.3 in transit; AWS service traffic over VPC endpoints (no public-internet traversal); LLM bodies not logged (token counts + content hashes only).

Caveats

  • This is a customer-architecture post on the AWS Architecture Blog (MHK, a Hearst Health company), not an AWS-internal retrospective. Numbers (90%, 85%, ~2 weeks) are MHK-reported business outcomes, not independently benchmarked latency/throughput.
  • No absolute throughput, p99 latency, concurrent-job counts, token-cost figures, or model-version specifics are disclosed beyond "hundreds of concurrent jobs" and "Claude".
  • MHK uses Bedrock exclusively for model inference and built its own orchestration, retrieval, and validation layers because healthcare needs domain-specific controls — i.e. this is deliberately not using Bedrock Agents / general-purpose guardrails for the workflow brain.
  • Future work (not yet shipped): conversational interfaces over the same workflow memory; growing the dynamic agent registry; exploring AWS Marketplace distribution to other regulated industries.

Taxonomy note

Source

Last updated · 766 distilled / 2,225 read