Google DeepMind Releases EmbeddingGemma 2, an Open Multimodal Embedder for Devices

Google DeepMind's open 740M-parameter EmbeddingGemma 2 maps text, code, images, audio and video into one embedding space and runs fully on-device.

October 6, 2026

Google DeepMind has launched EmbeddingGemma 2, a follow-up to last year’s text-only EmbeddingGemma, which has passed 20 million downloads. The new model is built on the Gemma 4 architecture and released under the Apache 2.0 license. It has 740 million parameters and maps text, code, images, video and audio into a single shared embedding space. That makes it possible, for example, to find a video clip from a voice memo or search hours of audio with a text query.

The model is modular: text-only use needs as little as 270M parameters, with optional vision (170M) and audio (300M) encoders. Matryoshka Representation Learning lets developers shrink vectors from 768 dimensions to 512, 256 or 128, cutting storage by up to 6x. With quantization on a Google Pixel 11 Pro, it needs about 191MB of active RAM for text-only and about 567MB for the full multimodal version. Its 8K-token context handles up to 5.5 minutes of audio, 29 images or 58 video frames. Google says it posts leading scores among sub-1B multimodal embedders on MTEB Code and MAEB, and MTEB Code rose from 68.76 to 78.68 versus the first version. Weights are on Hugging Face and Kaggle, with support across tools such as transformers.js, llama.cpp, Ollama, LMStudio, vLLM and MLX.

Why it matters

  • Privacy and offline use: Embeddings are generated locally, so search and retrieval across personal photos, recordings and files can work without sending data to the cloud and without an internet connection.
  • Cheaper for developers and businesses: The Apache 2.0 license is commercially permissive, and the small memory footprint and compressible vectors reduce storage and hardware needs for apps with local vector databases.
  • Better search across media and code: One model can handle text, images, audio, video and code, which could power features like finding video moments or indexing a local codebase for coding agents.

Source

Summary written by FoxaMind with AI assistance from the source above. Check the original for full details.