Google

How Gboard's New Private Learning System Handles Typing Data

Gboard now uses an auditable protected-server learning system for English and Japanese next-word prediction, with encrypted uploads, policy-bound decryption and explicit TEE limits.

Erawish 5 min read
Conceptual smartphone keyboard sending encrypted training data into sealed, auditable compute chambers
Conceptual smartphone keyboard sending encrypted training data into sealed, auditable compute chambers

Direct answer: Google says Gboard has deployed a new trusted-execution-environment-based federated-learning system for English and Japanese next-word prediction. Participating devices encrypt training examples before upload; only pre-authorized programs running inside attested protected servers can decrypt them, and those programs release differentially private model weights and metrics. This is a stronger and more auditable design, not a claim that typing-derived training data never leaves the phone or that protected hardware removes every risk.

What changed in Gboard's federated learning system?

Google Research announced the system on October 2, 2026. Earlier federated-learning designs generally computed model updates on participating devices and sent protected updates for aggregation. Google's new design lets devices upload encrypted training examples so much more of the training work can happen on servers inside trusted execution environments, or TEEs.

The practical change is not simply “more work happens in the cloud.” A participating phone encrypts each training upload and binds it to an access policy before sending it. That policy defines which server programs are allowed to process the upload. Google's key-management service releases decryption keys only to TEE workloads that match the approved policy.

Does Gboard training data leave the phone?

Yes, for the participating workloads described by Google, encrypted training examples leave the device. That is an important distinction from the shorthand claim that federated learning means all raw training material always stays on the phone.

Google says the uploaded examples remain encrypted outside the approved TEEs. The permitted program can decrypt and process them only inside those protected environments and only for a limited period after upload. The announcement does not provide ordinary users with a per-example dashboard, a list of every uploaded item or a universal promise that no metadata is visible.

Who can decrypt Gboard's encrypted training examples?

Google describes a chain of trust connecting the phone, a TEE-hosted key-management service and the later training workload. The key-management service gives a decryption key only to server binaries and Python training programs named in the upload's access policy.

Those access policies are published to Rekor, a public transparency log. Devices can require the policy to appear in that log before uploading, while external reviewers can inspect the set of programs that devices may authorize. Google also says the key-management and processing binaries can be reproduced from published open-source code.

Workload operators do not receive the decrypted examples directly under this design. Google says they can see released metrics and differentially private model weights instead.

What does differential privacy add?

Encryption and differential privacy solve different problems. Encryption and the TEE access policy limit who or what can open the training inputs. Differential privacy limits what the released model updates and metrics can reveal about an individual contribution.

In the accompanying paper, Google reports that a Gboard English-language experiment used 3.5 million devices in each test arm. The TEE-trained model used a three-times smaller reported privacy budget than the adjusted production baseline while remaining neutral on the paper's named typing metrics. It trained in about three weeks instead of two months. These are Google-reported experimental results, not an independent audit of every privacy or usability claim.

Which Gboard models use the new system?

Google says Gboard has deployed the system for English and Japanese next-word prediction models. The announcement does not say that every Gboard language, suggestion feature, keyboard setting or device now uses this exact pipeline.

This infrastructure story is also separate from Gboard's user-facing scam warnings. Erawish's September Pixel Drop guide covers the announced inline scam-warning feature and its regional limits; the new research announcement explains how selected prediction models are trained.

What privacy limits still remain?

  • Protected hardware is still a trust boundary: Google explicitly conditions the design on current-generation TEE limitations and points to ongoing work on side-channel risks.
  • Encrypted does not mean on-device only: selected training examples are uploaded in encrypted form for approved server-side processing.
  • Auditability has boundaries: published policies expose privacy-relevant programs, while proprietary model information can still be loaded at runtime when the hardcoded policy is designed to preserve the stated guarantees.
  • The announcement is not a settings guide: Google does not identify a new consumer-facing Gboard switch, universal rollout screen or per-upload inspection tool in this post.
  • The work is not presented as complete: Google calls the system a step toward rigorous proof and says deeper hardware protection and fuller correctness proofs remain future work.

What should Gboard users take away?

The main improvement is verifiability. Instead of asking users and researchers to trust an ordinary server operator's private handling of uploaded training material, Google is trying to make the allowed programs, protected execution path and anonymized outputs externally inspectable.

The careful reading is equally important: Gboard's new design sends encrypted training examples to protected servers, rather than promising that every training example remains on-device. Its privacy case depends on encryption, attestation, published access policies, differential privacy and the security of the TEE hardware working together.

Frequently asked questions

Does Google say it can read uploaded Gboard training examples outside the protected server?

Google says encrypted examples can be decrypted only inside TEE workloads that match the access policy bound to the upload. The company does not claim that all associated metadata is invisible or that TEEs eliminate every possible hardware or side-channel risk.

Is this the same as Gboard processing every keystroke in the cloud?

No. The announcement describes selected federated-learning training examples and specific English and Japanese next-word prediction models. It does not say that ordinary live typing or every keyboard action is sent to a server for immediate processing.

Can independent researchers verify the system?

Google says outsiders can inspect published access policies, transparency-log entries and reproducibly buildable open-source binaries. The accompanying paper still describes current TEE and side-channel limitations, so inspectability should not be confused with a claim that the entire system is risk-free or formally proven.

Sources