Skip to content

AWS 2026-09-18

Read original ↗

How CSIRO built scalable, cost-optimized genomic variant querying on AWS

Summary

CSIRO (Australia's national science agency), with the AWS ASP Prototyping and Scaling team, built Serverless Beacon (sBeacon) — a production implementation of the GA4GH Beacon protocol for securely querying genomic variant data. It is built entirely on AWS serverless primitives (Lambda, S3, DynamoDB, Athena), deployed per-institution via Terraform. The defining design choice is reference-not-copy: genomic VCF data is never ingested or copied; sBeacon builds index files that let queries read only the tabix-indexed region of interest directly from wherever the VCF lives (potentially a bucket owned by another organization) via HTTP byte-range requests. Variant queries use a fan-out / fan-in pattern of Lambda functions (Initiator → splitQuery → performQuery). The result is mega-biobank-scale capacity at ~USD 0.40/month for a 1000-Genomes-scale dataset, ~5 s query latency, and a zero-trust security posture with RBAC disclosure tiers enforced in the Lambda layer.

Key takeaways

  • Reference-not-copy indexing decouples cost from data volume. sBeacon does not copy the large genomic payload; it builds index files enabling random access, so storage cost is dominated by the small copied metadata (~1 MB of compressed ORC for 2504 samples). Ingesting chr1 (2504 individuals) took 18 seconds for < 1 cent (USD 0.00052). (Source: sources/2026-09-18-aws-how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws)
  • Two decoupled processes: onboarding and querying. Onboarding writes metadata to S3 in ORC, builds an ontology index via CSIRO Ontoserver, then runs Athena CREATE TABLE AS SELECT (CTAS) to build queryable metadata tables. Querying routes through API Gateway → a Microservice Lambda that resolves ontology terms in DynamoDB, queries metadata tables on Athena, and — if needed — invokes the Variant Querying Module.
  • The Variant Querying Module is a Lambda scatter-gather. An Initiator Lambda fans out splitQuery across the VCF files, which fans out performQuery across VCF regions within each file. performQuery fetches from S3 and returns results synchronously up the tree. Chose Lambda over Step Functions because Lambda processes much larger payloads for genomic fan-in/fan-out. (Source: sources/2026-09-18-aws-how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws)
  • Query latency is near-constant in result size. Querying a 10,000-base-pair region across 2504 individuals returns in ~1.52 s for USD 0.00013; latency stayed flat (1.51–1.65 s) as variants returned grew from 4 to 400.
  • Decentralization lives at the storage layer, not compute. A dataset's _vcfLocations are S3 URIs that can point to buckets owned by different organizations. performQuery passes each URI to bcftools, and htslib issues Range: bytes=X-Y requests (~1 KB per query) against that org's S3 REST API. Raw VCF bytes never leave the owning org; only the aggregate result (exists boolean / count / variant record) is returned upstream.
  • Zero-trust with RBAC disclosure tiers. Every request needs a valid JWT from an Amazon Cognito user pool, validated by the API Gateway authorizer before any Lambda runs. Authorization (disclosure granularity) is enforced inside the Lambda layer via Cognito group membership mapped to tiers: boolean (exists: true/false) → count → full record → admin. A boolean-tier user's request never causes sample-level data to be computed. (Source: sources/2026-09-18-aws-how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws)
  • Ephemeral compute isolation is a security property. Each Lambda cold start is a fresh container; /tmp (1,024 MB for performQuery) is cleared between cold starts; the bcftools subprocess runs and exits within the 10 s Lambda lifetime. No state persists after invocation (stateless compute).
  • Fan-out has a concurrency-exhaustion failure mode. A single fan-out query spawns many parallel Lambda invocations; bursts can exhaust the account concurrency pool and starve other functions (thundering herd on the concurrency budget). Synchronous Lambda invoke does not retry on throttle, so a 429 from performQuery silently loses that result unless the app handles it. Mitigation: alarm on Throttles/ConcurrentExecutions, request concurrency-limit increases, and enable provisioned concurrency on query-path Lambdas to cut cold-start latency during burst fan-out (at the cost of higher idle spend).
  • Whole-genome range queries time out at API Gateway. The architecture is tuned for population-scale region queries; an accidental whole-genome range query hits the API Gateway timeout. Genome-wide functional operations need further architectural work.

Systems / concepts / patterns extracted

Operational numbers

Case study: chromosome 1 (chr1, ~8% of the genome) of the 1000 Genomes Project, 2504 samples, multi-sample VCF ~1.1 GB compressed. All costs ap-southeast-2 (Sydney), pricing at time of writing.

  • Ingestion: chr1, 2504 individuals → 18 s, USD 0.00052.
  • Idle/maintenance: USD 0.000025/month (1 MB compressed metadata in ORC + genomic index files). Storing the genomic data too: USD 0.032 for chr1 (USD 0.025/GB), ~USD 0.425 for the whole genome.
  • Query: 10,000-bp region across 2504 individuals → 1.52 s, USD 0.00013; latency ~constant (1.51–1.65 s) as variants returned grow 4 → 400.
  • Whole-chr1 costs (per 1000 ops/month): ingestion USD 0.53 (32.82 GB-s Lambda); query compute USD 0.28 (9.8 GB-s Lambda); query Athena USD 0.05; idle storage (1.1 GB) USD 0.03; query DynamoDB USD 0.0005.
  • End-to-end generation-to-use lifecycle: ~18 s. Real-world query response: ~5 s.
  • performQuery limits: 1,024 MB /tmp, 10 s timeout; ~1 KB read per query via htslib byte-range.
  • Full-cohort claim: scales to hundreds of millions of individuals / billions of genomic locations (mega-biobank scale, per the sBeacon publication).

Caveats

  • This is an AWS Architecture Blog reference/solution post co-authored by CSIRO; cost/latency numbers are from a single published case study (chr1, 1000 Genomes) in one region, not fleet-wide production telemetry across many institutions.
  • The whole-genome-range-query timeout is an explicit known limitation.
  • Silent result loss on performQuery 429 throttles is a correctness caveat the operator must handle explicitly (no automatic retry on synchronous invoke).
  • Reference-not-copy assumes the source VCF stays available at its S3 URI; federated queries depend on cross-org bucket availability and read permissions.

Source

Last updated · 766 distilled / 2,225 read