Skip to content

SYSTEM Cited by 1 source

sBeacon (Serverless Beacon)

sBeacon (Serverless Beacon) is a production, open-source implementation of the GA4GH Beacon protocol — the Global Alliance for Genomics and Health's standard API for securely discovering genomic and phenotypic data across research and clinical networks — built entirely on AWS serverless services by CSIRO (Australia's national science agency) with the AWS ASP Prototyping and Scaling team. It is deployed per-institution as a container with Terraform-defined resources; each institution runs the full stack in its own AWS account, so there is no shared infrastructure, central data lake, or cross-account trust. (Source: sources/2026-09-18-aws-how-csiro-built-scalable-cost-optimized-genomic-variant-querying-on-aws)

Architecture

Two decoupled processes:

Data onboarding (indexing genomic metadata, not copying the payload):

  1. User submits the location of genomic data as request payloads to API Gateway.
  2. A Lambda handles indexing.
  3. Metadata written to S3 in ORC (columnar) format.
  4. An orchestrator Lambda drives the indexing process.
  5. CSIRO Ontoserver builds the ontology index for advanced metadata queries (also supports the Ensembl OLS V4 schema).
  6. Index files written to S3.
  7. Athena CREATE TABLE AS SELECT (CTAS) builds metadata tables. 8–9. Athena loads metadata from S3 and writes tables back to S3 in ORC.

Data querying (modular, split into Lambda functions by query scope):

  1. User query → API Gateway → Microservice Lambda.
  2. Microservice Lambda resolves query ontology terms in DynamoDB (descendant terms + codes).
  3. It queries metadata tables on Athena using those codes.
  4. If required, it invokes the Variant Querying Module.
  5. Result formatted per the Beacon protocol and returned via API Gateway.

Variant Querying Module — a serverless scatter-gather

The variant path is a Lambda fan-out / fan-in:

  • Initiator Lambda fans out splitQuery across the VCF files.
  • splitQuery fans out performQuery across VCF regions within each file.
  • performQuery fetches the VCF from S3 via bcftools/htslib byte-range reads and returns results synchronously up to the Initiator.
  • Optionally the Initiator also queries metadata from Athena (external table over S3).

CSIRO chose Lambda over Step Functions because Lambda processes much larger payloads, which fits the size/complexity of genomic data and the fan-in/fan-out shape.

Reference-not-copy: the core cost lever

sBeacon does not copy the large genomic payload. It builds index files that enable random access, so cost is driven by storing the small copied metadata. A dataset's _vcfLocations are S3 URIs that may point to buckets owned by entirely different organizations; performQuery passes each URI to bcftools, and htslib issues Range: bytes=X-Y requests (~1 KB per query) against that org's S3 REST API. Raw VCF bytes never leave the owning org's bucket — only the aggregate result (exists boolean / call_count integer / variant record) flows upstream. Decentralization is achieved at the storage layer, not the compute layer. This also keeps egress and data movement minimal.

Security — zero-trust with RBAC disclosure tiers

  • Explicit authN/authZ: every request needs a valid JWT from an Amazon Cognito user pool (COGNITO_USER_POOLS authorizer), validated by API Gateway before any Lambda runs (signature, expiry, audience). Auth can be disabled at first deploy (BEACON_ENABLE_AUTH = false) for intentionally public beacons — an explicit operator decision, not a default.
  • Least-privilege disclosure tiers (least privilege): Cognito groups map to disclosure granularity — sbeacon-boolean-access (exists true/false) → sbeacon-count-access (aggregate counts) → sbeacon-record-access (full variant + sample names) → sbeacon-admin (+ dataset management). The JWT carries group claims; the query Lambda reads them to set requested_granularity/include_details and passes both to performQuery, which computes only what was requested — a boolean-tier request never materializes sample-level data.
  • Ephemeral compute isolation (stateless): each cold start is a fresh container; /tmp (1,024 MB for performQuery) is cleared between cold starts; the bcftools subprocess runs and exits within the 10 s Lambda lifetime.
  • Cloud-native boundary controls: API Gateway is the only public entry; S3/DynamoDB/Athena/SNS have no public resource policies; Lambdas run in AWS-managed VPCs with no inbound network access.
  • Onboarding is not self-service: submitDataset requires sbeacon-admin membership beyond authentication.

Failure modes & tuning

  • Concurrency exhaustion (thundering herd on the concurrency pool): a single fan-out spawns many parallel invocations; monitor ConcurrentExecutions account- and function-level.
  • Silent result loss: synchronous Lambda invoke does not retry on throttle; a 429 from performQuery drops that result unless handled. Alarm on Throttles; a companion error-catcher architecture emails diagnostics.
  • Cold starts (cold start): after idle periods, simultaneous performQuery invocations hit cold starts; provisioned concurrency on query-path Lambdas trades idle cost for lower burst latency.
  • Whole-genome range query times out at API Gateway; genome-wide functional operations need further architecture.

Numbers (chr1, 1000 Genomes, 2504 samples, ap-southeast-2)

  • Ingest chr1: 18 s, USD 0.00052. Idle: USD 0.000025/month (~1 MB ORC metadata + index). Whole-genome-scale run: ~USD 0.40/month.
  • Query 10,000-bp region across 2504 individuals: ~1.52 s, USD 0.00013; latency ~constant as results grow 4 → 400. Real-world query ~5 s.
  • Scales to hundreds of millions of individuals / billions of locations (mega-biobank scale).

Seen in

Last updated · 766 distilled / 2,225 read