Skip to content

SYSTEM Cited by 1 source

WhatsApp Scam Alert

What it is

WhatsApp Scam Alert is an optional, user-controlled WhatsApp feature that runs an on-device machine-learning model to warn a user when an incoming message from a non-contact matches known scam patterns. The warning is shown only in the receiving user's chat and is not visible to the other party; from it the user can block, report, or continue — or mark the chat as trusted to suppress future warnings. No message content leaves the device for classification and nothing is auto-reported — reporting is always an explicit user action. Introduced in limited Beta (2026-08). (Source: sources/2026-08-12-meta-how-were-building-scam-alert-on-whatsapp-with-end-to-end-encryption-and-verifiability-guarantees)

The technically interesting part is not the classifier but the three architecturally-enforced, independently-verifiable guarantees built around it so it can coexist with WhatsApp end-to-end encryption. It is the third-generation instance of Meta's recurring privacy design — after Private Processing (private AI inference) and E2EE-backup key transparency — of pairing cryptography with hardware isolation and making the result publicly verifiable.

The three guarantees

# Guarantee Enforced by
1 On-device processing + privacy-preserving analytics — only DP-noised, k-anonymous aggregate counts ever leave the device on-device inference (concepts/on-device-ml-inference) + confidential federated analytics (federated-analytics) on TEEs (concepts/trusted-execution-environment) + DP (concepts/differential-privacy)
2 No targeted model delivery — every model version is published to a third-party append-only ledger before serving transparency ledger (model-transparency-ledger) + OHTTP (oblivious-http) + anonymous credentials (systems/meta-acs-anonymous-credentials) + non-targetability (concepts/defense-in-depth)
3 Verifiable model behavior — published weights + on-device transparency logs let researchers/users confirm scam-only purpose published model weights + Bug Bounty + in-app Scam Alert Activity log

Confidential federated analytics pipeline

The only telemetry that leaves the device is two aggregate categories: warning counts (measures precision / catches regressions across model versions) and user-action counts (trust vs. block-and-report; measures false-positive rate). The pipeline enforces:

  1. On-device data minimization — client aggregates raw signals into counts locally; metrics sent at randomized times, no device identifiers, coarse timestamps.
  2. Job selection — at randomized idle intervals under a daily resource cap, the client connects via OHTTP, authenticates with anonymous credentials, fetches active jobs, and rejects any job whose privacy parameters (ε, δ, k-anonymity) fail local guardrails.
  3. Attestation + RA-TLS — the client establishes an RA-TLS session with a stateless orchestrator TEE, cross-checking the attestation quote against the third-party transparency ledger before releasing data.
  4. Aggregator TEE — secure aggregation — merges metrics into running histograms, enforces k-anonymity suppression, applies DP noise, and bounds the ε/δ budget across all releases. Only noisy, thresholded aggregates cross the TEE boundary.
  5. Encrypted recovery checkpoints — long-running aggregates survive crashes; checkpoint keys never leave the TEE.

The client fails closed on the privacy invariant: if the TEE binary doesn't match the ledger or privacy parameters are insufficient, it refuses to transmit. The pipeline extends Meta's peer-reviewed PAPAYA Federated Analytics Stack (USENIX NSDI 2025).

Model delivery + verification

The model is served from a CDN (not baked into the app) so it can improve without a forced app upgrade. Delivery is designed for non-targetability:

  • Publication — server computes SHA-256 of weights/tokenizers/assets, builds a JSON manifest; the manifest digest is signed by a third-party signer (Cloudflare) using Ed25519 (Meta does not hold the signing key), published to a third-party append-only ledger, and assets uploaded to the CDN.
  • Anonymous download — via OHTTP (IP-stripped) + anonymous credentials; the request carries no identity selectors.
  • Client-side verification — verify Ed25519 signature against hardcoded Cloudflare public keys → cross-reference the manifest digest against the ledger with freshness (anti-replay) checks → verify each asset's SHA-256 against the manifest. Any failure → refuse to load.
  • Private experimentation — the client self-assigns its experiment group with a local random seed; group properties are immutable post-publication, sizes can only expand, and a minimum group size prevents targeting. Every variant is on the ledger; too-small groups are k-anonymity-suppressed.

See client-side-model-verification-against-ledger and client-side-experiment-assignment.

Threat model + defense-in-depth

Three attacker classes and their controls:

  • External actors intercepting data in transit → device↔TEE encryption; a compromised OHTTP relay sees only ciphertext + relay IP.
  • Insiders with infra access → TEEs prohibit remote shell access even from the host; software built only from checked-in source with multi-engineer change control; unaggregated data never readable outside the TEE.
  • Physical/remote TEE interference → encrypted DRAM, CVM hardening, host monitoring, OHTTP routing. Because TEE guarantees are not absolute, a successful targeted attack would require compromising the entire system in a way that is publicly discoverable. Canonical instance of cryptography-plus-tee-defense-in-depth.

Relationship to other Meta privacy systems

  • systems/whatsapp-private-processing — the direct architectural ancestor (TEE + attestation + OHTTP + anonymous credentials for private AI inference). Scam Alert reuses the same primitives for analytics rather than inference.
  • systems/whatsapp — host product; Scam Alert joins its security/privacy cluster.
  • Contrast with Google's Confidential Federated Analytics — a very similar four-layer (crypto + TEE + key-distribution + DP) design in the Android/Google ecosystem.

Seen in

Last updated · 766 distilled / 2,225 read