On October 6, 2026, Google DeepMind announced EmbeddingGemma 2 — a 740-million-parameter open-weights multimodal embedding model that converts text, code, images, audio, and video into numeric vectors in a single shared space of 768 dimensions. The model is based on the Gemma 4 architecture and is distributed under the Apache 2.0 license, which permits commercial use — according to Google, this makes it optimal for on-device inference.

In the official blog announcement the company lists five key features of the new model: leading quality for its size, modular design, memory efficiency, on-device optimization, and an extended context window.

Benchmark Leadership

EmbeddingGemma 2 is built on technology from the Gemini Embedding models and, according to Google, delivers the best quality for its size among multimodal embedding models with fewer than one billion parameters — particularly on the MTEB Code and MAEB benchmarks. Most notable is the improvement on code: the MTEB Code score rose from 68.76 to 78.68, an increase of 9.92 percentage points. Google notes this can be used for indexing local codebases, semantic code search, and retrieval for coding agents.

The model also set a new standard for sub-billion-parameter models on images, video, documents, and audio, in some cases outperforming even some narrowly specialized models — more than twice its size. Google said it has published full evaluation results and model data in the EmbeddingGemma 2 model card.

The previous version — EmbeddingGemma — was released last year as a lightweight text embedding solution for direct ranking, search, and linking of information on consumer hardware. According to Google, the developer community's response exceeded expectations: the model was downloaded more than 20 million times and used in on-device search tools and privacy-first RAG pipelines. The second version builds on exactly that foundation, adding code, images, video, and audio.

EmbeddingGemma 2 uses the same text tokenizer and audio encoder as Gemma 4, so running the two models together in one pipeline reduces total memory consumption. Google says that combined with generative models — such as Gemma 4 — it can power on-device RAG pipelines that understand complex multimodal data.

On-Device Performance and Privacy

The model has a modular structure: text-only operation requires just 270 million parameters, while full multimodal mode adds a 170-million-parameter vision encoder and a 300-million-parameter audio encoder. With quantization, on a Pixel 11 Pro the text weights consume about 191 MB of RAM when active; for the full multimodal model the figure is about 567 MB.

According to Google DeepMind, EmbeddingGemma 2 makes it possible to build search, reranking, and retrieval systems that run on everyday devices without a cloud connection.

Generating embeddings locally helps ensure data privacy, reduces latency, and opens the door to fully offline cross-modal search systems.

The context window is 8,000 tokens — four times larger than the first version. According to Google, such a window fits up to 5.5 minutes of audio, 29 images, or 58 video frames, all processed on local hardware. As practical examples, the company cites finding a needed video clip from a voice note or searching through hours of audio recordings with a text query.

Adjustable Vector Dimensions and Memory Savings

The Matryoshka Representation Learning technique lets developers dynamically shrink output vectors from 768 dimensions to 512, 256, or 128. This delivers up to 6x memory savings in local vector databases.

Ecosystem and Availability

Model weights are available on Hugging Face and Kaggle; tools such as Transformers, sentence-transformers, vLLM, llama.cpp, MLX, Ollama, and SGLang are supported. Qdrant compatibility is listed for storing vectors, and an Unsloth guide on fine-tuning has been published.

For on-device deployment, Google offers two paths on the Google AI Edge platform: MediaPipe for ready-made embedding, search, and decision tasks, and LiteRT for custom integration. Running in the browser via transformers.js or WebGPU is also supported. The LiteRT community page on Hugging Face hosts device-optimized models. Google said it has published a developer guide, documentation, and instructions on inference and fine-tuning. Google also said the AI Edge Gallery app will soon gain "Instant Media Search" and "Video Moments Finder" features. The Google DeepMind model page describes EmbeddingGemma 2 as designed for search, reranking, and retrieval systems that run on everyday devices without cloud dependence.