Overview
- Google DeepMind released EmbeddingGemma 2, which launched Tuesday, October 6, 2026, as an open-weight 740 million-parameter embedding model published under an Apache 2.0 license.
- The model natively maps text, code, images, video, and audio into a single 768-dimensional vector space and expands the context window to 8,192 tokens to cover longer audio and multi-frame inputs.
- EmbeddingGemma 2 is modular with a 270M-parameter text/code backbone and optional vision (~170M) and audio (~300M) encoders that together reach 740M parameters, and Google reports quantized RAM use of about 191MB for text-only and 567MB for the full multimodal setup on a Pixel 11 Pro.
- It uses Matryoshka Representation Learning so developers can truncate embeddings from 768 to 512, 256, or 128 dimensions for up to roughly sixfold storage savings, and Google reports strong benchmark gains especially on code retrieval.
- Weights and deployment guidance are available now on Hugging Face and Kaggle with tooling for LiteRT, MediaPipe, and browser runtimes, Google has published device demos such as AI Edge Foresight, and wider real-world tests will determine how index scale and truncation affect retrieval quality and storage on consumer devices.