Overview
- DeepSeek officially released V4.1 Flash on Thursday, Sept. 10, and made the model available through its API with published model weights and a technical report.
- V4.1 Flash is a 552B-parameter Mixture-of-Experts model using a Causal-Encoder-Decoder layout that DeepSeek says reduces activation sizes and compresses KV cache by large factors to lower runtime memory and storage needs.
- The new Flash unifies prior Fast/Expert/Vision modes into a single intelligent mode that auto-detects task complexity and activates vision capability when images are present.
- DeepSeek cut Flash-series API prices effective Sept. 10 and said V4 Pro requests will be routed to V4.1 Flash during a staged decommissioning of V4 Pro planned for Sept. 14.
- DeepSeek published the model and papers, reported partner integrations with Tencent and OpenCode, and said the architecture and KV-cache savings should reduce per-call costs for developers and agent-style workloads.