Particle.news

Google DeepMind Releases EmbeddingGemma 2, a 740M-Parameter Open Multimodal Embedder

The model lets developers run unified multimodal search on consumer hardware without sending private data to the cloud.

Overview

  • Google published EmbeddingGemma 2 on October 6 as an open‑weight model under an Apache 2.0 license with immediate downloads on Hugging Face and Kaggle.
  • The model maps text, code, images, audio and video into a single 768‑dimensional vector space so a text query can retrieve matching images, audio clips or code locally.
  • EmbeddingGemma 2 is modular: a 270M-parameter text/code base can be extended with a 170M vision encoder and a 300M audio encoder to reach 740M parameters so apps load only the parts they need.
  • Google reports practical on‑device footprints from Pixel 11 Pro tests: about 191MB active RAM for text-only and about 567MB for the full multimodal build, plus an 8K token context window for longer audio and video inputs.
  • The model uses Matryoshka Representation Learning so developers can truncate 768‑dimensional embeddings to 512, 256 or 128 dimensions to shrink local indexes, and Google notes strong sub‑1B benchmark gains especially for code retrieval.