Skip to content

META 2026-08-12

Read original ↗

How We're Building Scam Alert on WhatsApp With End-to-End Encryption and Verifiability Guarantees

Summary

Meta describes the architecture behind Scam Alert, an optional WhatsApp feature (early Beta, 2026-08) that runs an on-device machine-learning classifier to warn users about likely scam messages from non-contacts — without any message content leaving the device for classification and without auto-reporting. The interesting systems content is not the classifier itself but the three architecturally-enforced, independently-verifiable guarantees Meta builds around it so the feature can coexist with WhatsApp's end-to-end encryption: (1) on-device processing with privacy-preserving analytics — the only telemetry that leaves the device is aggregate warning and user-action counts, routed through a confidential federated analytics pipeline built on TEEs/CVMs and released only as differentially-private, k-anonymous aggregates; (2) no targeted model delivery — every model version (including experiment variants) is published to a third-party append-only transparency ledger and hash-verified on-device before use, with downloads routed through an OHTTP relay and anonymous credentials so no one can steer a specific model to a specific user; (3) verifiable model behavior — Meta publishes model weights and exposes in-app transparency logs so researchers and users can confirm the model is purpose-built for scams. The post is the third-generation instance (after WhatsApp Private Processing and Meta's E2EE-backup work) of Meta's recurring privacy pattern: pair cryptography with hardware isolation and make the whole thing publicly verifiable rather than merely promised.

Key takeaways

  1. The classifier is on-device by construction, not by policy. Scam Alert downloads a small ML model to the device and runs inference locally over incoming messages from non-contacts; the warning is shown only to the receiving user and is not visible to the sender. No message content is sent to WhatsApp/Meta for classification, and nothing is auto-reported — reporting happens only on explicit user action, consistent with existing WhatsApp user-reporting. This is a concrete instance of on-device ML inference used as a privacy boundary (Source: this source).

  2. "Is the feature working?" is answered with differentially-private aggregates, never per-user telemetry. Only two categories of approximate, aggregate counts leave the device: warning counts (how often the model surfaced a warning — measures precision / catches regressions across model versions) and user-action counts (trust vs. block-and-report — measures false-positive rate). Both are processed inside a confidential federated analytics pipeline and released only as differentially-private aggregates above a minimum cohort size (Source: this source).

  3. The analytics pipeline is a confidential federated analytics stack built on CVMs with client-verified attestation. Client aggregates raw signals locally into counts; metrics are sent at randomized times, carry no device identifiers, and coarsen timestamps. The client establishes an RA-TLS session with a stateless orchestrator TEE, cross-checking the attestation quote against a third-party transparency ledger before releasing data; an aggregator TEE merges into running histograms, enforces k-anonymity thresholds, and applies differential-privacy noise under a bounded ε/δ privacy budget across all releases. Only noisy, thresholded aggregates cross the TEE boundary. The pipeline extends Meta's peer-reviewed PAPAYA federated-analytics stack (USENIX NSDI 2025) (Source: this source).

  4. The client "fails closed" if privacy invariants aren't met. Before transmitting, the client verifies (a) the TEE binary matches the ledger-published measurement and (b) the job's privacy parameters (DP ε/δ, k-anonymity thresholds) meet locally-enforced guardrails. If verification fails or parameters are insufficient, the client refuses to transmit — any attempt to weaken the guarantees either fails closed or becomes publicly discoverable. This is a device-side fail-closed posture on the privacy invariant, not just availability (Source: this source).

  5. No targeted model delivery: every model version is published to a third-party append-only ledger before it is served to anyone. The model is served from a CDN (not baked into the app), so it can improve without a forced app upgrade — but every version's SHA-256 hashes are recorded in a signed manifest on a public, tamper-evident ledger before delivery. Delivering a targeted model would require publishing it on a ledger anyone can inspect, making the attempt publicly discoverable. Users can audit their own model via https://akd-auditor.cloudflare.com/namespaces/<namespace>/audits/<epoch> using the namespace + epoch from their in-app transparency log (Source: this source).

  6. Meta does not hold the model-signing key; Cloudflare (a third party) does. The manifest digest is signed by a third-party signer (Cloudflare) using Ed25519 keys, then published to the ledger; assets go to the CDN. The client verifies the signature against hardcoded Cloudflare Ed25519 public keys — confirming the manifest was signed by Cloudflare, not forged by Meta. This separation of the signing authority from the model author is the structural root of the non-targetability claim (Source: this source).

  7. Model download is anonymous and non-targetable. All download requests authenticate with anonymous credentials and route through an OHTTP relay that strips the requester's IP; the request payload contains no identity selectors, and the CDN serves only publicly-published files. Client-side verification is multi-step: verify the Ed25519 signature → cross-reference the manifest digest against the ledger with freshness checks (anti-replay) → verify each asset's SHA-256 against the manifest. If any step fails, the client refuses to load the model (Source: this source).

  8. Private experimentation is done via on-device group assignment, so experiments can't become a targeting side-channel. Using a locally-generated random seed, the client assigns itself to an experiment group and selects its model variant from the download response — the server cannot steer a specific user to a specific variant. Integrity checks enforce that experiment group properties can't change after publication, experiment sizes can only expand (never shrink), and groups must meet a minimum size threshold (preventing an attacker from narrowing a group to target an individual). Every experiment variant is on the ledger; experiment metrics flow through the same confidential federated analytics pipeline, and groups too small to meet k-anonymity are suppressed entirely (Source: this source).

  9. Defense-in-depth against a three-category threat model. Meta models three attacker classes — third-party/supply-chain vendors, malicious/compromised insiders, and external actors — and layers controls accordingly: encryption between device and TEE (a compromised OHTTP relay sees only ciphertext + relay IP), TEEs that prohibit remote shell access even from the host, software built only from checked-in source with multi-engineer change control, and encrypted DRAM + CVM hardening + host monitoring for physical/remote TEE interference. Because TEE guarantees are not absolute, a successful targeted attack would require compromising the entire system in a way that is publicly discoverable through verifiable transparency (Source: this source).

  10. Verifiable model behavior closes the "which binary, doing what" gap. Attestation proves which model is running, not what it does — so Meta publishes model weights (for AI/ML researchers, via an expanded Bug Bounty program) and exposes on-device transparency logs (Account > Request Info > Scam Alert Activity) recording which messages were flagged, the outcome, and the model version used. Encrypted recovery checkpoints let the long-running aggregation survive crashes without restarting, with keys that never leave the TEE (Source: this source).

Systems / concepts / patterns extracted

Systems

Concepts

Patterns

  • on-device-classify-aggregate-privately (new) — classify locally, emit only DP/k-anonymous aggregate counts through a confidential pipeline.
  • client-side-model-verification-against-ledger (new) — client verifies signature + ledger membership + per-asset hash before loading a model.
  • client-side-experiment-assignment (new) — device self-assigns experiment groups with locally-generated randomness so the server can't target.
  • cryptography-plus-tee-defense-in-depth — the umbrella composition pattern this system instantiates.

Operational details / numbers

  • Stage: limited Beta rollout (2026-08); an "early technical overview" published ahead of general availability.
  • Telemetry surface: exactly two aggregate metric categories — warning counts, user-action counts (trust / block-and-report).
  • Signing: Ed25519, third-party signer (Cloudflare); Meta does not hold the signing key. Model asset integrity via SHA-256 hashes in a signed manifest.
  • Privacy parameters: differential-privacy ε/δ + k-anonymity thresholds, checked client-side against local guardrails; DP budget bounded across all releases.
  • Routing: OHTTP relay (IP-stripping) + anonymous credentials for both metrics upload and model download.
  • User audit path: in-app Account > Request Info > Scam Alert Activity; ledger audit URL https://akd-auditor.cloudflare.com/namespaces/<namespace>/audits/<epoch>.
  • Foundations: confidential federated analytics pipeline extends Meta's PAPAYA Federated Analytics Stack (USENIX NSDI 2025).

Caveats

  • Early Beta, not GA — the design may evolve; some capabilities ("we will be publishing…", "we will provide in-app capabilities…") are stated as forthcoming.
  • TEE guarantees are explicitly not absolute — Meta relies on defense-in-depth and public discoverability of tampering rather than claiming the TEE alone is unbreakable; see tee-side-channel-vulnerability.
  • Attestation ≠ correctness — publishing weights + transparency logs is what substantiates the "purpose-built for scams" claim; attestation alone only proves which binary runs.
  • The post is a security-architecture overview, not a full engineering white paper (Meta commits to publishing the latter for the analytics pipeline).

Source

Last updated · 766 distilled / 2,225 read