How Clario technology detects PHI/PII in DICOM images using Amazon Bedrock¶
Summary¶
Clario (part of Thermo Fisher Scientific) built a production AWS pipeline that automatically detects PHI (Protected Health Information) and PII in DICOM medical-imaging files across clinical trials, where a single imaging series can span thousands of slices and sensitive data hides in three surfaces: standard DICOM header tags, vendor-specific custom private tags, and text physically burned into the image pixels. The detection workflow is exposed through Amazon API Gateway (TLS, IAM-backed auth, API-key validation, rate limiting), runs its long-running, memory-intensive backend on Amazon EKS, uses Amazon Textract for OCR/structure extraction and Anthropic's Claude Sonnet 4.5 on Amazon Bedrock for PHI/PII classification (including a deep pixel scan), and persists processing metadata in PostgreSQL on Amazon RDS for a compliance audit trail. The defining design decision is separating detection from redaction: the pipeline returns bounding-box coordinates and PHI types to downstream QC and redaction systems, keeping a human-in-the-loop reviewer in control of the irreversible act of modifying clinical data.
Key takeaways¶
- Three detection surfaces, one pipeline. PHI/PII can live in (1) standard DICOM metadata tags, (2) non-standard vendor/site-specific private tags, and (3) burned-in pixel text. Measured detection F1 / label accuracy: DICOM metadata tags 0.9951 / 99.60%, PDF text 0.9775 / 98.12%, DICOM burned-in image text 0.9750 / 96.15%. The solution scanned 100% of image slices, standard headers, and custom private tags in the internal test dataset. (Source: sources/2026-08-19-aws-how-clario-detects-phi-pii-in-dicom-images-using-bedrock)
- EKS chosen for long-running, memory-intensive work. Because a single DICOM series can span thousands of slices, the detection workload is long-running and memory-heavy, so the backend runs on Amazon EKS rather than a short-lived function runtime; AWS SA guidance covered pod-scaling strategy to handle thousands of slices per request without timeout or memory pressure. (Source: sources/2026-08-19-aws-how-clario-detects-phi-pii-in-dicom-images-using-bedrock)
- RDS PostgreSQL for the audit trail. Processing metadata is persisted in PostgreSQL on Amazon RDS specifically because the compliance audit trail needs relational queries and strong consistency for compliance reporting.
- API Gateway as the edge choke point. The service is fronted by API Gateway so authentication, API-key management, and rate limiting are handled at the edge — keeping the EKS backend focused purely on detection (see gateway-strips-auth-dispatches-cached-backend).
- Textract → Bedrock, not off-the-shelf detectors. Many open-source and off-the-shelf detection models looked acceptable on curated samples but degraded badly on production variability (diverse modalities, vendor private tags, inconsistent burned-in text). The winning pipeline is a tuned composition: Textract handles text extraction, Claude Sonnet on Bedrock performs classification, with prompt engineering tuned to clinical-trial DICOM patterns.
- Deep pixel scan for burned-in PHI. Beyond metadata, Claude Sonnet 4.5 on Bedrock performs a deep scan of actual image pixel content to catch patient names, DOBs, and IDs physically rendered into every slice — returning spatial coordinates and PHI type for each detection (see visual-first-document-extraction and concepts/multi-modal-attribute-extraction).
- Detection separated from redaction (the headline decision). The pipeline only identifies PHI (coordinates + type) and hands results to two downstream flows — a QC flow where qualified reviewers validate flagged findings, and a redaction flow that masks burned-in pixel text or strips/zeroes sensitive DICOM tags. This preserves flexibility, keeps auditability clear, and keeps a human in the loop before any irreversible change.
- Ground truth is non-negotiable. A trustworthy automated evaluation required hand-annotated ground truth (exact PHI text + bounding boxes). Matching strategy: spatial proximity with a default 3-pixel tolerance per bounding-box element where coordinates exist; structured-path matching for coordinate-less DICOM metadata. (see ground-truth-from-analyst-feedback)
- Data-minimization via TTL. Ingested files and records are retained only for a limited window: the S3 document is auto-deleted by an S3 Lifecycle policy and RDS records by a scheduled cleanup job, to meet HIPAA/GDPR data-minimization requirements (see patterns/ttl-based-resource-lifecycle).
- Cost tuning at scale. AWS SA collaboration recommended batching Textract API calls and optimizing Claude Sonnet prompt-token usage to reduce per-document inference cost, and tuned concurrent Bedrock invocations to maximize throughput within account-level quotas.
Architecture¶
Data flow (Figure 1 in the post):
Consumer (upstream imaging app)
│ 1. upload DICOM (.dcm) or PDF
▼
[[Amazon S3]] (source bucket)
│ 2. call Clario Internal API with API key + document location
▼
[Amazon API Gateway](<../systems/amazon-api-gateway.md>) ← TLS, IAM auth, API-key validation, rate limiting
│ 3. validate key → forward to backend
▼
Detection backend on [[Amazon EKS]] ← long-running, memory-intensive
│ 4. initial checks (URL format, access, basic metadata)
│ 5. retrieve file, copy to internal S3, persist metadata to RDS
▼
[Amazon Textract](<../systems/amazon-textract.md>) ← 6. OCR + tables + form fields → structured text
│
▼
[Amazon Bedrock](<../systems/amazon-bedrock.md>) (Claude Sonnet 4.5) ← 7. classify PHI/PII in OCR text
│ + deep pixel scan of image slices
│ 8. return structured coordinates + PHI type
▼
Consumer Downstream QC flow (human review)
Downstream redaction flow (mask pixels / strip tags)
│ 9. TTL cleanup
▼
S3 Lifecycle auto-delete + RDS scheduled cleanup job
Metadata storage: PostgreSQL on Amazon RDS (relational queries + strong consistency for the audit trail). Observability across logs, metrics, and traces via Amazon CloudWatch and AWS CloudTrail; security via IAM, VPC controls, and encryption in transit and at rest, all inside Clario-managed AWS accounts.
Operational numbers¶
| Detection surface | Detection F1 | Label accuracy |
|---|---|---|
| DICOM metadata tags | 0.9951 | 99.60% |
| PDF text | 0.9775 | 98.12% |
| DICOM burned-in image text | 0.9750 | 96.15% |
- Coverage: 100% of image slices, standard DICOM headers, and custom private tags scanned in the internal test dataset.
- Ground-truth spatial match tolerance: 3 pixels per bounding-box element (default).
- Scale driver: a single DICOM series can span thousands of slices; supported file
types are
.dcmand PDF. - Clario context (external): >50 years in endpoint solutions, deployed >30,000 times, supporting >700 FDA/EMA new-drug approvals since 2015.
Lessons learned (from the post)¶
- Evaluate models against production-representative data — off-the-shelf detectors degrade on real-world variability; a tuned Textract + Claude Sonnet pipeline outperformed them.
- Ground-truth data is non-negotiable — manual annotation is expensive but prerequisite to a trustworthy automated evaluation.
- Separating detection from masking improves flexibility and auditability.
- Human-in-the-loop review remains an essential safeguard for edge cases and model uncertainty in a regulated domain.
Caveats¶
- Metrics are from internal testing on a Clario-assembled dataset, not a large-scale external benchmark; no throughput (docs/sec), latency, cost-per- document, EKS pod counts, or Bedrock concurrency numbers are disclosed.
- The redaction/QC systems are described only as downstream consumers of the detection output; their internals are out of scope.
- Prompt-engineering specifics, Textract batching parameters, and RDS/S3 retention windows are described qualitatively, not quantified.
- This is an AWS Architecture Blog write-up co-authored with the Clario team, so it carries some reference-architecture / solution-marketing framing alongside the real production details.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/how-clario-automates-phi-pii-detection-in-dicom-images-using-amazon-bedrock/
- Raw markdown:
raw/aws/2026-08-19-how-clario-technology-detects-phipii-in-dicom-images-using-a-14b93ff5.md
Related¶
- managed-ai-document-intelligence-pipeline-on-aws — the reusable AWS document-intelligence composition this instantiates
- concepts/human-in-the-loop — detection surfaces uncertainty; a human decides before irreversible redaction
- systems/amazon-textract · systems/amazon-bedrock · systems/aws-eks · systems/aws-rds · systems/amazon-api-gateway
- visual-first-document-extraction — pixel-level PHI detection over burned-in text