Particle.news

Google Releases EmbeddingGemma 2, a 740M Open Multimodal On‑Device Embedder

The model promises low‑memory multimodal search that keeps user data on devices rather than sending it to the cloud.

Overview

  • Google launched EmbeddingGemma 2 on Tuesday, October 6 and published the model weights under an Apache 2.0 license on Hugging Face and Kaggle for immediate developer use.
  • The model is modular and totals 740 million parameters with a 270M text/code backbone plus optional 170M vision and 300M audio encoders that all map into a single 768‑dimensional embedding space.
  • EmbeddingGemma 2 expands context to 8,192 tokens and uses Matryoshka Representation Learning to let developers truncate vectors from 768 to 512, 256, or 128 dimensions to cut local index storage up to about six times.
  • Google reports quantized on‑device footprints of roughly 191 MB RAM for a text‑only setup and about 567 MB for the full multimodal configuration on a Pixel 11 Pro, and it showcased demos such as AI Edge Foresight, Instant Media Search, and Video Moments Finder.
  • Google shares benchmark claims including a 9.92‑point gain on MTEB Code to 78.68, but independent evaluations and real‑world tests of large local indexes, long‑term adoption, and quality tradeoffs at smaller vector sizes remain to be seen while developers begin integrating the model.