How CSIRO built scalable, cost-optimized genomic variant querying on AWS¶
Summary¶
CSIRO (Australia's national science agency), with the AWS ASP Prototyping and
Scaling team, built Serverless Beacon (sBeacon) — a production
implementation of the GA4GH Beacon protocol for securely querying genomic
variant data. It is built entirely on AWS serverless primitives (Lambda,
S3, DynamoDB, Athena),
deployed per-institution via Terraform. The defining
design choice is reference-not-copy: genomic VCF data is never ingested or
copied; sBeacon builds index files that let queries read only the tabix-indexed
region of interest directly from wherever the VCF lives (potentially a bucket
owned by another organization) via HTTP byte-range requests. Variant queries
use a fan-out / fan-in pattern of Lambda
functions (Initiator → splitQuery → performQuery). The result is
mega-biobank-scale capacity at ~USD 0.40/month for a 1000-Genomes-scale dataset,
~5 s query latency, and a zero-trust
security posture with RBAC disclosure tiers enforced in the Lambda layer.
Key takeaways¶
- Reference-not-copy indexing decouples cost from data volume. sBeacon does not copy the large genomic payload; it builds index files enabling random access, so storage cost is dominated by the small copied metadata (~1 MB of compressed ORC for 2504 samples). Ingesting chr1 (2504 individuals) took 18 seconds for < 1 cent (USD 0.00052). (Source: sources/2026-09-18-aws-how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws)
- Two decoupled processes: onboarding and querying. Onboarding writes
metadata to S3 in ORC, builds an ontology index via
CSIRO Ontoserver, then runs Athena
CREATE TABLE AS SELECT(CTAS) to build queryable metadata tables. Querying routes through API Gateway → a Microservice Lambda that resolves ontology terms in DynamoDB, queries metadata tables on Athena, and — if needed — invokes the Variant Querying Module. - The Variant Querying Module is a Lambda scatter-gather.
An
InitiatorLambda fans outsplitQueryacross the VCF files, which fans outperformQueryacross VCF regions within each file.performQueryfetches from S3 and returns results synchronously up the tree. Chose Lambda over Step Functions because Lambda processes much larger payloads for genomic fan-in/fan-out. (Source: sources/2026-09-18-aws-how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws) - Query latency is near-constant in result size. Querying a 10,000-base-pair region across 2504 individuals returns in ~1.52 s for USD 0.00013; latency stayed flat (1.51–1.65 s) as variants returned grew from 4 to 400.
- Decentralization lives at the storage layer, not compute. A dataset's
_vcfLocationsare S3 URIs that can point to buckets owned by different organizations.performQuerypasses each URI to bcftools, and htslib issuesRange: bytes=X-Yrequests (~1 KB per query) against that org's S3 REST API. Raw VCF bytes never leave the owning org; only the aggregate result (exists boolean / count / variant record) is returned upstream. - Zero-trust with RBAC disclosure tiers. Every request needs a valid JWT
from an Amazon Cognito user pool, validated by the
API Gateway authorizer before any Lambda runs. Authorization (disclosure
granularity) is enforced inside the Lambda layer via Cognito group
membership mapped to tiers: boolean (
exists: true/false) → count → full record → admin. A boolean-tier user's request never causes sample-level data to be computed. (Source: sources/2026-09-18-aws-how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws) - Ephemeral compute isolation is a security property. Each Lambda cold start
is a fresh container;
/tmp(1,024 MB forperformQuery) is cleared between cold starts; thebcftoolssubprocess runs and exits within the 10 s Lambda lifetime. No state persists after invocation (stateless compute). - Fan-out has a concurrency-exhaustion failure mode. A single fan-out query
spawns many parallel Lambda invocations; bursts can exhaust the account
concurrency pool and starve other functions (thundering herd
on the concurrency budget). Synchronous Lambda invoke does not retry on
throttle, so a
429fromperformQuerysilently loses that result unless the app handles it. Mitigation: alarm onThrottles/ConcurrentExecutions, request concurrency-limit increases, and enable provisioned concurrency on query-path Lambdas to cut cold-start latency during burst fan-out (at the cost of higher idle spend). - Whole-genome range queries time out at API Gateway. The architecture is tuned for population-scale region queries; an accidental whole-genome range query hits the API Gateway timeout. Genome-wide functional operations need further architectural work.
Systems / concepts / patterns extracted¶
- Systems: sBeacon (subject), bcftools / htslib, CSIRO Ontoserver, AWS Lambda, Amazon S3, Amazon DynamoDB, Amazon Athena, Amazon API Gateway, Amazon Cognito, Terraform.
- Concepts: serverless compute, scale-to-zero, scatter-gather query (Lambda fan-out/fan-in), zero-trust authorization, least-privilege access, stateless / ephemeral compute, columnar storage format (ORC metadata), cold start, thundering herd (concurrency exhaustion), egress cost (byte-range reads keep bytes in the owning account).
- Patterns: none minted — the fan-out/fan-in variant query is an instance of scatter-gather over serverless compute; reference-not-copy indexing, tabix byte-range reads, and RBAC disclosure tiers are recorded as tags + prose (single-source, article-specific framings).
Operational numbers¶
Case study: chromosome 1 (chr1, ~8% of the genome) of the 1000 Genomes Project, 2504 samples, multi-sample VCF ~1.1 GB compressed. All costs ap-southeast-2 (Sydney), pricing at time of writing.
- Ingestion: chr1, 2504 individuals → 18 s, USD 0.00052.
- Idle/maintenance: USD 0.000025/month (1 MB compressed metadata in ORC + genomic index files). Storing the genomic data too: USD 0.032 for chr1 (USD 0.025/GB), ~USD 0.425 for the whole genome.
- Query: 10,000-bp region across 2504 individuals → 1.52 s, USD 0.00013; latency ~constant (1.51–1.65 s) as variants returned grow 4 → 400.
- Whole-chr1 costs (per 1000 ops/month): ingestion USD 0.53 (32.82 GB-s Lambda); query compute USD 0.28 (9.8 GB-s Lambda); query Athena USD 0.05; idle storage (1.1 GB) USD 0.03; query DynamoDB USD 0.0005.
- End-to-end generation-to-use lifecycle: ~18 s. Real-world query response: ~5 s.
performQuerylimits: 1,024 MB/tmp, 10 s timeout; ~1 KB read per query via htslib byte-range.- Full-cohort claim: scales to hundreds of millions of individuals / billions of genomic locations (mega-biobank scale, per the sBeacon publication).
Caveats¶
- This is an AWS Architecture Blog reference/solution post co-authored by CSIRO; cost/latency numbers are from a single published case study (chr1, 1000 Genomes) in one region, not fleet-wide production telemetry across many institutions.
- The whole-genome-range-query timeout is an explicit known limitation.
- Silent result loss on
performQuery429throttles is a correctness caveat the operator must handle explicitly (no automatic retry on synchronous invoke). - Reference-not-copy assumes the source VCF stays available at its S3 URI; federated queries depend on cross-org bucket availability and read permissions.
Source¶
- Original: https://aws.amazon.com/blogs/architecture/how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws/
- Raw markdown:
raw/aws/2026-09-18-how-csiro-built-scalable-cost-optimized-genomic-variant-quer-f3494b00.md
Related¶
- systems/sbeacon — the subject system
- systems/bcftools-htslib — the genomic byte-range read tool
- systems/csiro-ontoserver — ontology terminology server for metadata queries
- concepts/scatter-gather-query — the fan-out/fan-in variant query
- concepts/serverless-compute, concepts/scale-to-zero — the cost model
- concepts/zero-trust-authorization — the security posture
- companies/aws