Overview
- DeepSeek disclosed Aug. 13 that DeepSeek‑V4‑Pro‑0813 is now generally available across its API, app and web, released under an MIT license as an open‑weight Mixture‑of‑Experts model with a 1 million‑token context window.
- The startup is switching to peak and off‑peak billing effective Aug. 16, raising V4‑Pro output to $3.96 per 1M tokens at peak (half that off‑peak) and setting V4‑Flash peak output at $1.32 per 1M tokens.
- DeepSeek published internal benchmark tables showing the 0813 build close to Anthropic’s Claude Fable 5 on several agent tests, but independent third‑party verification of the new weights is still limited.
- Rapid token growth has strained capacity on an estimated ~20,000 H100 GPU footprint, so DeepSeek is accelerating hiring, fundraising and plans for more chips and data‑center capacity to manage heavier, always‑on agent workloads.
- The company also released an MIT‑licensed DeepSeek Harness to help developers build agentic coding stacks and the open weights that let firms fine‑tune models privately, a step that eases adoption but raises data‑sovereignty and regulatory questions for some buyers.