Particle.news

DeepSeek’s V4‑Flash Benchmarked as Cheapest‑to‑Run Major AI Model

This cost reduction could reshape the economics of deploying AI for businesses, cloud providers and decentralized compute networks.

Overview

  • Independent benchmarking by Artificial Analysis shows DeepSeek’s V4‑Flash charges about $0.14 per million input tokens and $0.28 per million output tokens and runs at roughly three cents per test, making it the least expensive of well‑known models.
  • The research firm found V4‑Flash more than 100 times cheaper than Anthropic’s Claude Fable 5 on some tests while scoring 50 out of 100 on its Intelligence Index, a midrange result that signals tradeoffs between cost and top‑tier capability.
  • DeepSeek attributes the low inference cost to a Mixture‑of‑Experts architecture and hardware/software co‑design that activate fewer parameters per request, and prior reporting says the firm trained an earlier V3 line for roughly $5.6 million in GPU rentals.
  • Markets and rivals are reacting: DeepSeek is reported to be preparing a larger V4‑Pro and exploring an IPO, incumbents like Alibaba unveiled bigger models, and cheaper inference is changing demand for cloud GPUs and the business case for DePIN and other decentralized compute projects.
  • Lower per‑query costs and DeepSeek’s open MIT license make AI deployment far cheaper for small businesses and developers, which could broaden real‑world use cases, reduce total ownership costs for AI products, and alter who captures value in the AI stack.