CONCEPT Cited by 2 sources
Federated learning¶
Definition¶
Federated learning (FL) is a machine-learning setting where multiple entities (clients — typically end-user devices) collaborate to solve an ML problem under the coordination of a service provider, without centralizing the raw training data. Instead of shipping user data to a central server, the training algorithm is brought to the data: clients compute on their local data and only privacy-reduced artifacts (model-gradient updates, or — in newer designs — encrypted examples processed inside attested enclaves) leave the device.
Google's evolved (2025) definition centers FL on four privacy principles: (1) data minimization, (2) data anonymization, (3) transparency and control, and (4) verifiability and auditability. A complete FL system should let clients retain full control over their data, the set of workloads allowed to access it, and the anonymization properties of those workloads. (Source: sources/2026-10-02-google-toward-provably-private-learning-from-federated-data)
Why it exists¶
The canonical motivating problem (from Google's 2017 FL work) is improving an on-device model — e.g. Gboard's next-word-prediction language model — from real user behavior without ever shipping keystrokes off-device. FL provides the mechanism: local model updates aggregated across a fleet, with anonymization layers (secure aggregation, differential privacy) bounding what the aggregate can reveal about any one user. Production FL at Google has powered Gboard next-word prediction and Smart Compose, Google Messages reply suggestions, and Android Smart Text Selection. (Source: sources/2026-10-02-google-toward-provably-private-learning-from-federated-data)
Generations of FL architecture¶
The anonymization/trust model has evolved through distinct generations (Source: sources/2026-10-02-google-toward-provably-private-learning-from-federated-data):
| Generation | Where compute runs | Trust model | Weakness addressed next |
|---|---|---|---|
| Gen 1 (2017) client-side aggregation | Gradients computed on-device; server aggregates | Server could log/inspect uploads; no external verification | No verifiability |
| + Secure Aggregation | On-device, cryptographically masked uploads | Server sees only the sum | Incompatible with state-of-the-art central DP (MF-DP-FTRL) |
| Gen (TEE-based, this post) server-side verifiable training | Devices upload encrypted examples; the full training loop runs server-side inside attested TEEs | Operator trust removed via attestation + transparency log; only DP weights leave the enclave | Current-gen TEE / side-channel limits; formal proofs still future work |
The Gen-TEE design (see systems/confidential-federated-compute) is a notable inversion: it moves gradient computation off the device and onto the server, trading the device-side trust model for a hardware-rooted, externally verifiable server-side one. This also removes on-device-compute and diurnal device-availability bottlenecks, cutting Gboard model training from 1–2 months to TEE-availability-limited. (Source: sources/2026-10-02-google-toward-provably-private-learning-from-federated-data)
Relationship to federated analytics¶
Federated analytics is the same decentralized-data, server-coordinated pattern applied to analytics/aggregate statistics rather than model training. Google's Confidential Federated Analytics is the aggregation-direction sibling of the FL training system; both compose TEEs + attestation + transparency logs over the same device fleet. (Source: sources/2026-05-27-google-private-analytics-via-zero-trust-aggregation)
What FL is not¶
- Not inherently private. Raw gradients can leak training data; FL requires explicit anonymization layers (secure aggregation and/or DP) to bound leakage.
- Not "no server." A service provider still coordinates rounds, aggregates, and (in the TEE generation) runs the training loop; FL constrains what the server can see, not whether a server exists.
- Not free of the training/serving boundary — FL redraws where training compute happens but the serving model still ships to devices.
Seen in¶
- sources/2026-10-02-google-toward-provably-private-learning-from-federated-data — the TEE-based, server-side, verifiable FL generation; Gboard adoption.
- sources/2026-05-27-google-private-analytics-via-zero-trust-aggregation — the federated-analytics sibling (zero-trust aggregation).
Related¶
- systems/confidential-federated-compute — the Gen-TEE FL system
- systems/google-confidential-federated-analytics — analytics-direction sibling
- systems/gboard — canonical production FL target
- concepts/trusted-execution-environment — trust boundary for the new generation
- concepts/differential-privacy — the anonymization layer on released weights
- concepts/on-device-ml-inference — the device-side serving counterpart
- concepts/training-serving-boundary — FL redraws where training runs
- companies/google