Overview
- DeepSeek has moved V4‑Flash into public beta and released a benchmark claiming its per-task AI cost is about 60% lower than GPT-5.6 Luna.
- OpenRouter, a major model-API aggregator, reported V4‑Flash led weekly global API token calls with 7.22 trillion tokens and OpenCode posted an 8 trillion-token single-day spike on its platform.
- The heavy reported usage is already driving industry responses, with major AI vendors lowering product prices and strengthening entry-level model offerings.
- Key figures and cost comparisons come from DeepSeek and platform reports, and industry observers note those numbers require independent third-party verification.
- If sustained, V4‑Flash’s low reported costs and large scale could cut AI costs for developers, shift provider choice toward lower‑cost models, and pressure mid-tier vendors’ market positions.