Toward provably private learning from federated data¶
Summary¶
Google Research announces the next generation of its Federated Learning (FL) system, which replaces trust in the server operator with externally verifiable, hardware-rooted privacy guarantees built on Trusted Execution Environments (TEEs). Rather than aggregating client gradients on-device (the 2017-era FL design), the new system has devices encrypt and upload training examples, then runs the entire training loop server-side inside attested TEEs that can only execute computations the device pre-authorized via a published access policy. The architecture coordinates four parts — encrypted data upload under an access policy, a RAFT-based Key Management System (KMS) that releases decryption keys only to attested workloads matching the policy, a root data-processing TEE that delegates parallel subtasks to a cluster of worker TEEs, and KMS-encrypted recovery state for fault tolerance. Access policies are published to Rekor (a public transparency log) and binaries are reproducibly built from the open-source Confidential Federated Compute repository, so external auditors can verify exactly what code touches user data. Shifting compute to the server removed on-device-compute and diurnal-availability bottlenecks: Gboard next-word-prediction models that previously took 1–2 months to train now train far faster, limited only by TEE resource availability.
Key takeaways¶
-
Server-side FL with verifiable privacy replaces "trust the operator." Earlier FL uploaded data for immediate aggregation with no way for outside observers to verify it was never logged or inspected; later Secure Aggregation added cryptographic protection but was incompatible with state-of-the-art central differential privacy. The new TEE-based system is the "next milestone in the ongoing effort to completely remove the need to trust the server operator." (Source)
-
Four coordinated operational concepts. (1) Data upload: devices locally encrypt training examples and pre-authorize an access policy (the set of TEE computations allowed to process the data, which only release anonymized results); access policies must be published to a public transparency log. (2) KMS and policy verification: the KMS, a cluster of TEEs running the RAFT consensus protocol, releases decryption keys only to server-side TEE workloads whose measurement matches the access policy. (3) Workload execution: a "root" data-processing TEE runs a Python training loop and delegates parallelizable subtasks to a cluster of worker TEEs, periodically releasing anonymized (DP) model weights. (4) Fault-tolerant recovery: the program saves KMS-encrypted recovery state at the end of each round to recover from root/worker failures without leaking additional privacy-sensitive information. (Source)
-
Only DP model weights and metrics ever leave the TEE boundary. Encrypted device data can be decrypted and processed only inside TEEs running the Python training programs named in the access policies, and only for a limited time after upload. Workload operators see metrics and differentially private weights, never raw data. (Source)
-
Public transparency log + reproducible builds make the system auditable. Access policies describing allowed server workloads are published to Rekor (part of Sigstore); external auditors watch the log to track the full set of workloads devices could participate in. The KMS and data-processing binaries are reproducibly built from the open-source Confidential Federated Compute GitHub repository, so the attested measurement can be tied back to source. (Source)
-
Dynamic sideloading protects proprietary logic while preserving auditability. Access policies published to Rekor directly describe the Python training program. To protect proprietary model architectures and preprocessing, data-processing TEEs support sideloading serialized information into the Python program at runtime — allowed as long as all privacy-relevant logic remains hardcoded in the published program. This keeps the externally verifiable privacy guarantee intact while letting non-privacy logic stay proprietary. (Source)
-
Distributed logic is written in Federated Language. The training loop is expressed in Federated Language, an open-source, framework-agnostic orchestration language derived from TensorFlow Federated (which powered the previous FL system). (Source)
-
Shifting compute to the server removed on-device and diurnal bottlenecks. Collecting all uploads before running the server-side workload eliminates diurnal device-availability variation as a training-progress constraint; the optimal device-participation schedule and other DP parameters can be computed dynamically at execution time. Training that used to take 1–2 months per model — limited by device availability, on-device compute, and cross-workload contention for device resources — is now substantially faster, bottlenecked only by TEE resource availability. (Source)
-
Gboard is the first production adopter. Gboard launched English and Japanese next-word-prediction models on the new system with stronger privacy guarantees and improved accuracy; it is "benefiting from substantially faster compute times than our previous FL system." (Source)
-
Enables larger FL models and arbitrary verifiable Python workloads. Moving client-gradient computation to the server lifts on-device compute limits, paving the way for training larger models via FL; integrating TEEs with accelerators is called out as important future work. The same infrastructure can run any workload expressible in Python verifiably — Google is experimenting with synthetic-data generation and combining these data-processing TEEs with LLM-inference TEEs. (Source)
Systems extracted¶
- Confidential Federated Compute — the new TEE-based FL system itself (root + worker TEEs, RAFT KMS, reproducibly-built open-source binaries). The GitHub repo of the same name is the reproducible-build source of truth.
- Google Confidential Federated Analytics — the sibling confidential-federated-analytics stack this work builds on (plus "provably private insights"); same TEE + attestation + transparency family, aggregation direction.
- Gboard — first production adopter (EN + JA next-word prediction).
- Sigstore / Rekor — the public append-only transparency log where access policies are published for external audit.
- Federated Language — open-source framework-agnostic orchestration language for distributed FL logic, derived from TensorFlow Federated.
Concepts extracted¶
- Federated learning — the overarching ML setting; this post defines its evolution from client-side aggregation to verifiable server-side training.
- TEE — the hardware trust boundary (remotely attestable, confidential, integrity-protected) that the whole system is built on.
- Remote attestation — how the KMS verifies a workload's measurement before releasing decryption keys.
- Differential privacy — only DP model weights are released from the training loop.
- Consensus / RAFT — the KMS is a cluster of TEEs running RAFT to agree on key release.
- Confidential computing — the umbrella for the "data in use" protection this provides.
- Side-channel attack — explicitly named as the residual TEE risk and a dynamic-sideloading caveat.
- Training/serving boundary — the shift of gradient computation from device to server redraws this boundary.
- Confidential storage inside TEE — KMS-encrypted recovery state for fault tolerance.
- Audit trail — Rekor transparency log + reproducible builds give external auditors a verifiable record.
Operational numbers & specifics¶
- Prior FL model training: 1–2 months per model, bottlenecked by device availability, on-device compute, and cross-workload device contention.
- New system: substantially faster, bottlenecked only by TEE resource availability; server-side parallelization across many worker TEE machines.
- KMS: a cluster of TEEs running RAFT; releases keys only to attested workloads matching the access policy.
- Decryption window: encrypted uploads decryptable only for a limited time after upload, and only inside policy-named Python training TEEs.
- First production models: English and Japanese Gboard next-word prediction.
- Repos:
google-parfait/confidential-federated-compute(reproducible build source),google-parfait/federated-language(orchestration language). - Transparency log: Rekor (
docs.sigstore.dev).
Caveats¶
- Subject to current-generation TEE limitations — the post explicitly cites TEE limitations and side-channel observations as residual risks; the verifiable-privacy guarantee is "subject to current-generation TEE limitations." Future hardware + side-channel-mitigation research is expected to deepen protection, especially for dynamically-loaded (sideloaded) workloads.
- Dynamic sideloading is a controlled escape hatch, not unconstrained. Its safety rests on the invariant that all privacy-relevant logic stays hardcoded in the published Python program; sideloaded serialized data must be privacy-irrelevant.
- Formal proofs are aspirational. The post is "a step toward rigorous proof" and anticipates that such systems "may one day come with full proofs of correctness of the software implementations of the DP algorithms and system components" — i.e. end-to-end formal verification is future work, not shipped.
- No ε/δ DP budgets, committee/cluster sizes, or aggregation-latency numbers are disclosed; the post is architecture-focused.
- Malicious-client / data-poisoning defenses are not discussed.
Source¶
- Original: https://research.google/blog/toward-provably-private-learning-from-federated-data/
- Raw markdown:
raw/google/2026-10-02-toward-provably-private-learning-from-federated-data-b9b18411.md