Particle.news

Google Reportedly Building Gemini-Specific Server Chip Called Frozen v2

To embed parts of the Gemini model into silicon, Google is testing an experimental chip that aims to cut inference energy use.

Overview

  • A report published Monday says Google is developing a server chip, internally called Frozen v2, that would hard‑wire parts of the Gemini model architecture into silicon to reduce on‑server computation and data movement.
  • Engineers quoted in the coverage estimate the design could process about six to ten times more tokens per unit of energy than Google’s most recent TPUs, a measure tied directly to lower cloud inference costs.
  • Google is reportedly planning a limited deployment around 2028 and frames Frozen v2 as a complement to its existing TPU line rather than a direct replacement.
  • The push responds to internal compute shortages that have strained Google Cloud capacity and led the company to buy extra capacity from outside vendors, and the effort aims to lessen reliance on dominant suppliers such as Nvidia.
  • The project is experimental and carries tradeoffs: tightly coupling hardware to Gemini raises risks if future model architectures change, investors bid Alphabet shares up roughly 3%, and internal delays on Gemini releases highlight organizational hurdles to turning the research into a mass product.