Overview
- DeepSeek announced a planned, substantial increase to its API pricing and told users to reasonably arrange their usage while final rates are not yet published.
- The company currently bills by tokens with separate charges for input tokens and output tokens and differentiates prices for cache hits versus cache misses.
- Published current rates show two main models: V4-Flash priced for cost and high concurrency, and V4-Pro priced for performance and complex reasoning.
- Cache misses are far more expensive than cache hits so applications with low cache-hit rates will see proportionally larger bill increases once new prices take effect.
- Customers are being asked to reassess model choice, tighten caching and usage patterns, and await DeepSeek’s formal notice for the definitive new rates and billing details.