Google Open-Sources EmbeddingGemma 2: Lightweight Multimodal Embedding Model for 0.5GB RAM On-Device AI
Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 740M multimodal embedding model mapping text, audio, and vision into a unified 768-dim vector space
On October 6, 2026, Google DeepMind officially unveiled EmbeddingGemma 2, a lightweight, open-source multimodal embedding model that projects text, code, images, audio, and video into a single unified 768-dimensional vector space. Released under the permissive Apache 2.0 license, the sub-1B parameter model totals 740M parameters and is engineered to run directly on consumer smartphones, laptops, and edge hardware without relying on cloud infrastructure.

Image source: Google DeepMind
While traditional embedding models have either focused strictly on text or required multi-gigabyte GPU memory pools on centralized cloud servers for multimodal tasks, EmbeddingGemma 2 delivers native multimodal representation inside a modular, sub-1B architecture optimized for on-device local RAG and semantic retrieval.
768-Dimensional Unified Vector Space and 740M Modular Architecture
The defining technical breakthrough of EmbeddingGemma 2 is its unified latent space: distinct media modalities are not segregated into separate vector embeddings, but are mapped directly into a shared 768-dimensional coordinate system.
Built upon the Gemma 4 decoder architecture, the model divides its 740M total parameters across a modular design:
- Text and Code Base Model (270M): Consists of a 130M parameter backbone and a 140M parameter embedder dedicated to handling natural language and code tokens.
- Vision Encoder (170M): Encodes static images and video frames directly into 768-dimensional embeddings.
- Audio Encoder (300M): Projects spoken audio memos and acoustic waveforms into the same vector space alongside text and visual embeddings.
This unified alignment enables native cross-modal retrieval out of the box. Users can search through video footage using spoken audio clips, or query hours of audio recordings using brief text prompts, all evaluated within a single vector distance metric.
0.5GB RAM On-Device Footprint and Privacy-First Local RAG
EmbeddingGemma 2 is tailored to run smoothly on client hardware, including standard laptops, mobile processors, and embedded systems without discrete GPUs.
When operating in text-only mode, the model requires approximately 0.5GB of RAM—and can shrink to roughly 191MB depending on runtime quantization formats like bfloat16 or GGUF. This compact footprint unlocks robust offline capabilities for privacy-sensitive environments:
- Fully Offline Local RAG: Index and search personal documents, private notes, and internal codebases on local storage without streaming data to third-party cloud servers.
- On-Device Semantic Search and Clustering: Automatically tag, organize, and categorize images, audio recordings, and text files on the client device even when disconnected from the internet.
- Zero Data Leakage: Protect sensitive enterprise intellectual property and confidential records by keeping all retrieval and embedding computations strictly on-device.
Apache 2.0 Licensing and Framework Integration Caveats
Google DeepMind has released EmbeddingGemma 2 under the Apache 2.0 license, permitting unrestricted commercial integration alongside academic research.
Model weights are immediately available on Hugging Face (google/embeddinggemma-2), with out-of-the-box support across Sentence Transformers and Unsloth for GGUF-quantized inference. Developers integrating the model into edge pipelines should note several critical configuration details:
- Default Multi-Encoder Loading: Standard loaders such as Sentence Transformers will load all 740M parameters across vision, audio, and text encoders by default.
- Text-Only Memory Optimization: For workflows requiring only text and code embedding, developers must explicitly disable auxiliary encoders by passing
config_kwargs={"vision_config": None, "audio_config": None}to achieve the compact 270M parameter profile (~0.5GB RAM). - Quantization and Latency Variability: Operating memory consumption and vector indexing latency will vary based on target CPU/NPU cache bandwidth and the chosen quantization level (bfloat16 versus GGUF).
Sources
- Google DeepMind / Google Blog: EmbeddingGemma 2: an open, lightweight multimodal embedding model
- Google Developers Blog: EmbeddingGemma 2: The Developer Guide
- Google AI for Developers: EmbeddingGemma 2 Model Card
- Hugging Face: google/embeddinggemma-2
- Hugging Models Official X (@HuggingModels): EmbeddingGemma 2 Release Announcement