Particle.news

Google Reportedly Developing Frozen v2 Chip to Run Gemini Models Far More Efficiently

If verified, the specialized chip could sharply lower energy costs for serving AI queries at scale.

Overview

  • Monday reports say Alphabet is building a server chip nicknamed Frozen v2 that would hardwire parts of its Gemini model into silicon to speed inference, though Google has not confirmed the project.
  • Sources quoted in the coverage claim Frozen v2 could serve six to ten times more AI tokens per watt than Google’s current custom processors, a gain that would far exceed the typical 2–3x per‑generation improvement.
  • The Information and other outlets report engineers are still finalizing the design and the company has discussed deploying the chip as soon as 2028 while one outlet offered an earlier timeline, leaving timing uncertain.
  • Alphabet is positioning Frozen v2 as a new, specialized family separate from its TPUs to optimize inference, control supply lines, and reduce reliance on third‑party GPUs such as Nvidia’s hardware.
  • Key questions remain about whether the lab efficiency claims will hold in real data centers, how much of Gemini will be fixed in silicon, and how any cost‑per‑inference gains would change cloud pricing and enterprise AI adoption.