Overview
- DeepSeek announced V4.1 Flash on Thursday, Sept. 10, describing a slimmed‑down multimodal model that it says uses a causal‑encoder‑decoder architecture and Mixture‑of‑Experts (MoE) routing to activate only small parameter subsets for each request.
- The company reported Terminal‑Bench 2.1 results that put V4.1 Flash at 90.6, above OpenAI’s GPT‑5.6 Sol (88.8), Moonshot AI’s Kimi K3 (88.3) and DeepSeek’s V4‑Pro (87.9), but those scores are published by the vendor and lack broad independent verification.
- DeepSeek says the 552‑billion‑parameter framework routes work so that inputs use about 8 billion active parameters and responses use about 16 billion, which the firm argues cuts latency and operating cost to as little as a fraction of a cent per million tokens.
- The company plans to retire V4‑Pro by automatically rerouting inference to V4.1 Flash starting Sept. 14, a move that coincided with share drops in some Hong Kong‑listed rivals and price pressure across the market.
- If third‑party tests confirm DeepSeek’s claims, enterprises and startups could run longer, cheaper agentic workflows such as code writing and image reasoning, but buyers must weigh the vendor‑reported gains against limited outside validation and rising hardware and export constraints.