RU

EmbeddingGemma 2: an open multimodal embedding model with 740 million parameters

Published: 2026-10-07 · Author: AI Release · @ai_release1
EmbeddingGemma 2: an open multimodal embedding model with 740 million parameters

⚡ The gist in 5 seconds - Google DeepMind has released EmbeddingGemma 2 — an open multimodal embedding model with 740 million parameters that unifies text, code, images, video, and audio in a single vector space. - The model is available on Hugging Face and Kaggle under the Apache 2.0 license and is optimized for local inference on consumer devices. - Limitation: full multimodality requires ~567MB of active RAM (with quantization on the Pixel 11 Pro), while text-only mode needs ~191MB. ### 🔍 What was found Google has announced EmbeddingGemma 2 — the second version of the family introduced a year earlier as a lightweight solution for text embeddings. The first model racked up over 20 million downloads and was used for local search and privacy-first RAG pipelines. The new version is built on the Gemma 4 architecture and extends capabilities beyond text: a single embedding space covers code, images, video, and audio. The authors are Google DeepMind engineers Sahil Dua and Enrique Schechter Vera. Key technical specs: the modular design requires as few as 270M parameters for text tasks, with optional encoders — vision (170M) and audio (300M). Thanks to Matryoshka Representation Learning, vectors can be truncated from 768 down to 512, 256, or 128 dimensions, yielding up to 6x storage savings for local vector databases. The context window is 8K tokens (4x larger than the first version), allowing local processing of up to 5.5 minutes of audio, 29 images, or 58 video frames. On the MTEB Code and MAEB benchmarks, the model leads among multimodal embedders under 1B parameters, posting a 9.92-point gain in code (from 68.76 to 78.68) and outperforming some specialized models twice its size. ### 💡 Why it matters Local embedding generation ensures data privacy, reduces latency, and enables cross-modal search and RAG that run fully offline. Since the model is built on Gemma 4 and shares its tokenizer and audio encoder, both models can be run in a single pipeline with lower total memory consumption — for example, EmbeddingGemma 2 for file search plus Gemma 4 for contextual reasoning. Practical scenarios are already available in Google AI Edge Gallery: Instant Media Search for searching your media library, Video Moments Finder for finding moments in video via text or audio queries, and the MediaPipe Decision Task API for real-time classification and routing. ### 🧩 Context The first EmbeddingGemma appeared a year earlier as a compact model for high-quality text embeddings on consumer hardware — for organizing, searching, and linking information directly on-device. The developer response exceeded Google's expectations: over 20 million downloads, with adoption in local search tools and privacy-first RAG. EmbeddingGemma 2 builds on this line, adding multimodality while keeping the bet on edge inference. The model is compatible with a broad tooling stack: transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGL

🔗 Read on blog.google

🤖 AI summary
GoogleEmbeddingGemmaэмбеддингиon-deviceоткрытые
← PreviousWindows 11 is testing a Windows XP-style window title listNext →Windows: One Command Solves the Problem — Not Enough Permissions to Modify a File

Source: blog.google · post in Telegram