Particle.news

Microsoft Unveils Maia 200, an Inference Accelerator Built Around Software-Controlled Dataflow

The company says Maia 200 pairs a Software-defined Local Access dataflow with an all‑Ethernet scale-up network to lower cost and energy per generated token.

Overview

  • Microsoft presented Maia 200 as its second-generation custom AI accelerator that treats software, memory, network, and silicon as a single co‑designed system to optimize inference.
  • Maia 200 centers on SDLA, a software-driven dataflow that gives programs explicit control of data movement between high‑bandwidth memory and on‑chip SRAM to produce steadier, more predictable kernel performance.
  • Microsoft reported internal benchmarks showing roughly 30% better performance-per-dollar than its latest GPUs and more than 40% higher token generation on the MAI-Thinking-1 model under the same rack power budget.
  • The design uses a two-tier, Ethernet-based HammingMesh scale-up network that Microsoft describes as able to span rack domains and scale to thousands of accelerators to make networking part of the execution engine.
  • Independent outlets published specific hardware numbers for Maia 200 such as TFLOPS, 750W TDP, and 7 TB/s HBM bandwidth but those figures are reported by third parties and have not been independently verified in the coverage provided.