Overview
- Microsoft presented Maia 200 as its second-generation custom AI accelerator that treats software, memory, network, and silicon as a single co‑designed system to optimize inference.
- Maia 200 centers on SDLA, a software-driven dataflow that gives programs explicit control of data movement between high‑bandwidth memory and on‑chip SRAM to produce steadier, more predictable kernel performance.
- Microsoft reported internal benchmarks showing roughly 30% better performance-per-dollar than its latest GPUs and more than 40% higher token generation on the MAI-Thinking-1 model under the same rack power budget.
- The design uses a two-tier, Ethernet-based HammingMesh scale-up network that Microsoft describes as able to span rack domains and scale to thousands of accelerators to make networking part of the execution engine.
- Independent outlets published specific hardware numbers for Maia 200 such as TFLOPS, 750W TDP, and 7 TB/s HBM bandwidth but those figures are reported by third parties and have not been independently verified in the coverage provided.