Google Releases EmbeddingGemma 2: Lightweight 740M Multimodal On-Device Embedding Model
Google has launched EmbeddingGemma 2, an open 740M-parameter multimodal embedding model for text, code, images, video, and audio on local devices, now on Huggin
On October 6, 2026, Google officially announced EmbeddingGemma 2, an open multimodal embedding model optimized for on-device efficiency, and immediately released its model weights on Hugging Face.

Image Credit: @sundarpichai / X
Marking Google's first open, natively multimodal embedding model, EmbeddingGemma 2 is built into a lightweight, modular 740M-parameter form factor capable of indexing diverse modalities within a unified model architecture.
Unified Multimodal Embeddings Across Five Data Modalities in 740M
The core technical milestone of EmbeddingGemma 2 is its native ability to process modalities beyond conventional text inside a compact, on-device model.
- Supported Modalities: Handles text, source code, images, video, and audio tasks within a single model architecture.
- Lightweight Modular Form Factor: Scaled to 740 million parameters, allowing efficient execution on resource-constrained local environments including mobile devices, laptops, and embedded edge systems.
- Outperforming Larger Baselines: According to Google, the model outperforms certain specialized embedding models more than twice its size across benchmark tasks.
This architecture enables developers to eliminate fragmented pipelines that previously required distinct modality-specific embedding models or costly proprietary cloud APIs.
Pairing with Gemma 4 for Privacy-First Offline RAG
EmbeddingGemma 2 is specifically optimized for pairing with Gemma 4 on-device generative models to power local Retrieval-Augmented Generation (RAG).
Developers can index and retrieve local documents, codebases, photos, and audio recordings without sending payloads to external cloud servers. This makes it an ideal foundation for privacy-first, fully offline AI workflows where sensitive corporate code or private personal data must remain confined to the local device.
Availability and Deployment Considerations
EmbeddingGemma 2 model weights are available immediately as open source via Hugging Face for integration into local pipelines and research frameworks.
As a 740M compact on-device model, users should note natural trade-offs compared to massive cloud-hosted embedding services regarding ultra-long context windows and complex multi-step retrieval tasks. Detailed token context limits and specific multilingual benchmark scores should be verified against the official model card documentation on Hugging Face.
Sources
- Sundar Pichai Official X (@sundarpichai): EmbeddingGemma 2 Announcement