EmbeddingGemma 2: What Runs Locally Now and What's Coming
EmbeddingGemma 2 can be downloaded today as a multimodal retrieval model and tried in current AI Edge demos; Android ML Kit service and Model Garden routes are still coming.
Direct answer: EmbeddingGemma 2 is Google's new 740-million-parameter model for turning text, code, images, video frames and audio into comparable numerical representations called embeddings. Developers can download the model now and try Google's current local-search demos. Google says an Android ML Kit service and Gemini Enterprise Agent Platform Model Garden availability are still coming, so this is not yet a built-in search feature on every Android or Pixel phone.
The practical value is that one compact model can help an app match a natural-language query against several kinds of local media. It is a retrieval component, not a chatbot: it helps find and rank related items, while another system must decide what to show or do with those results.
What is EmbeddingGemma 2?
Google launched EmbeddingGemma 2 on October 6, 2026 as an open-weight, Apache 2.0 model built for multimodal embeddings. Instead of generating an answer, it maps an input into a vector—a list of numbers that represents semantic meaning. Items with similar meaning should land near one another in that vector space.
The model accepts text, including source code, plus images, sampled video frames and audio. That shared representation can support tasks such as finding a photo from a written description, locating a moment in a video, matching an audio query to related media or indexing a local codebase for semantic search.
"Multimodal" does not mean the model understands every file or application automatically. The integrating app still has to select and prepare the input, create an index, run similarity search and decide how results are presented.
What can developers use today?
Google's launch materials identify several current ways to evaluate or build with the model:
- Model weights and documentation: Google links current downloads through Hugging Face and Kaggle, along with the model card, inference guidance and fine-tuning documentation.
- AI Edge Gallery demos: Instant Media Search can match text or an example image against selected local photos and videos. Video Moments Finder can index video segments and return timestamps that match a written query.
- AI Edge Foresight on Mac: Google's experimental desktop app uses local models to index meeting transcripts and private files for note and context retrieval.
- Direct developer routes: Google's implementation article documents MediaPipe Tasks and LiteRT paths, plus pre-quantized bundles and reference examples. Exact platform, accelerator and package support still needs to be checked for each deployment.
These are developer tools and demonstrations. They do not establish that Android Photos, Files, Recorder or another consumer app has adopted EmbeddingGemma 2, and they do not make the model a system feature on a particular phone.
What is still coming?
Two announced routes are not presented as generally available today:
- Android ML Kit service: Google says EmbeddingGemma 2 will become available as an Android service through ML Kit in the coming weeks, including NPU acceleration where supported.
- Gemini Enterprise Agent Platform Model Garden: Google's launch post lists this as coming soon rather than a current download route.
That current-versus-coming split matters. Developers can experiment with the downloadable model and today's AI Edge examples, but they should not write production plans as if the future ML Kit service contract, eligible-device coverage or rollout date were already final.
How much memory does EmbeddingGemma 2 use?
Google reports about 191 MB of active RAM for quantized text-only weights and about 567 MB for the full multimodal model on a Pixel 11 Pro. It also describes a modular design: a text core can be used alone, with vision and audio encoders added when an application needs those inputs.
Those figures are useful reference points, not universal minimum requirements. They come from Google's measurement on one named device and configuration. Real memory use, latency, battery impact and sustained performance will vary with quantization, enabled encoders, input size, runtime, accelerator, operating system and the rest of the application.
Google also publishes benchmark and latency results in its launch post and model card. Treat them as vendor-reported evaluations for the documented tasks, not proof that the model will outperform alternatives in every language, media library or production workload.
Does local processing guarantee privacy?
No. Running the embedding calculation and index on a device can reduce the need to upload selected media, and Google's current demos are designed to show offline local retrieval. But the model release cannot guarantee how every integrating app handles permissions, source files, vector indexes, analytics, logs, synchronization or generated results.
A privacy review still needs to ask where each processing step runs, which data the app can access, what is retained, whether any telemetry leaves the device and how a user can delete the source and derived index. "Uses an on-device model" is evidence about one component, not the whole application's data flow.
How is this different from AISeal and Gboard private learning?
These Google technologies solve different problems:
- Android AISeal is an architecture for hardware-isolated personal-context storage, with in-vault model execution and agents described as future work.
- Gboard's private-learning system concerns how selected encrypted training examples are processed and audited in server-side trusted execution environments.
- EmbeddingGemma 2 is a downloadable retrieval model that creates embeddings for local or edge search. It does not by itself provide a protected vault, application permissions or a complete training system.
Using the three names together should not become a claim that all Google AI processing is local, that every model runs inside AISeal or that an application receives privacy guarantees merely by adopting EmbeddingGemma 2.
What remains unknown?
- The exact release date, API contract and eligible-device coverage for the announced Android ML Kit service.
- When EmbeddingGemma 2 will appear in Gemini Enterprise Agent Platform Model Garden.
- How individual consumer applications will request access to personal media, store indexes and disclose local-versus-cloud processing.
- How performance changes across lower-memory phones, desktops, browsers and different CPU, GPU or NPU backends.
- How retrieval quality varies by language, media type, user library and application-specific data.
Bottom line
EmbeddingGemma 2 makes multimodal retrieval available as a relatively compact developer model today: weights, documentation and hands-on AI Edge examples are current. The Android ML Kit service and enterprise Model Garden route remain announced next steps. For readers evaluating it now, the right question is not simply whether it is "on-device," but which exact workload, runtime, data path and hardware have actually been tested.
Sources
- Google: EmbeddingGemma 2 launch, architecture, memory figures, downloads and availability boundaries
- Google Developers: AI Edge demos, MediaPipe, LiteRT and announced ML Kit path
- Google AI for Developers: EmbeddingGemma 2 model card
- SiliconANGLE: technical reporting on the launch and multimodal expansion