September 12, 2026 · DeepSeek
DeepSeek Releases V4.1 Flash with 1 Million Token Context at $0.15 per Million Input
My take: The pricing pressure from China keeps coming, and this is another chapter in the same story. DeepSeek released V4.1 Flash on September 10 at a base price of $0.15 per million input tokens and $0.60 per million output tokens, placing it among the cheapest models in its class. The architecture is a mixture of experts with 552 billion total parameters and 8 billion active per token, supporting up to one million tokens of context with text and image input. The most notable technical improvement is a 75% reduction in key-value cache memory usage per token, which translates to faster inference and lower operational cost for those running it locally.
The pattern that matters is not this model specifically but what it represents: the rates budgeted in August are no longer current market rates in September. Competition between models is not slowing down.
For any team using AI in production, the practical question is whether your contracts and budgets reflect current prices, or whether you have gone months without reviewing whether what you pay is still competitive. How often do you evaluate the alternatives available in the model market?
Want to use these tools? See the unbiased reviews or back to the news.