Particle.news

AMD Details MI455X and Puts Helios Servers Into Production as Token Demand Rises

The new accelerators promise much larger memory and throughput, plus vendors warn that network and operational capacity must scale to handle surging token workloads.

Overview

  • AMD on July 24 unveiled technical specs for the Instinct MI455X, saying the chip uses TSMC 2nm parts, carries about 320 billion transistors and pairs with 432 GB of HBM4 to raise per‑GPU memory and bandwidth.
  • The company also said its second‑generation Helios AI rack servers are in full production and expected to ship by the end of Q3 2026, with customers including major AI users preparing deployments.
  • Huawei reported token consumption rose roughly sixfold in six months and argued that network and operational throughput, not just single‑chip process gains, are now the primary bottleneck for large‑scale model use.
  • Reporting showed talent and morale shifts at major AI labs, including reported resignations at Tencent and departures from DeepMind that sources link to internal disputes and a controversial DoD agreement, factors that may slow flagship model timelines.
  • Automakers continue heavy product and operations pushes: FAW and Zeekr launched new PHEV and EV models this week while XPeng reiterated a plan for Robotaxi per‑vehicle breakeven by the second half of 2027, signaling parallel scaling and cost pressures in mobility.