Overview
- Google announced EmbeddingGemma 2 Tuesday as a 740 million-parameter model built on the Gemma 4 architecture that natively maps text, images, audio and video into a single embedding space and is released under an Apache 2.0 license.
- The model is modular so developers can load smaller encoder combinations for text-only or add vision and audio encoders, and it uses Matryoshka Representation Learning to let embeddings be truncated from 768 down to 512, 256 or 128 dimensions to cut storage by as much as sixfold.
- Google reports that with quantization EmbeddingGemma 2 can run with about 191 MB of active RAM for text-only and about 567 MB for the full multimodal configuration on a Pixel 11 Pro, targeting low-memory on-device use cases.
- Google says EmbeddingGemma 2 leads sub-1B benchmark rankings on tasks in MTEB and MAEB and shows a near 10-point gain on code tasks, though independent third-party verification of those claims is not provided in the coverage.
- Weights and tools are publicly available on Hugging Face and Kaggle and integrations include LiteRT, MediaPipe, transformers, transformers.js and llama.cpp, which Google and partners are using to demo apps such as an AI Edge notetaking app that processes transcripts and audio entirely on-device.