DeepSeek Flash Beats GPT-5 at a Fraction of the Cost
DeepSeek V4 Flash 0731 jumped from 56.9 to 82.7 on Terminal Bench and from 51.8 to 70.3 on Toolathlon in a single update, clearing GPT-5 Terra (78.4) on the coding benchmark. Two separate threads covered this, and the reaction was notably warmer than typical model announcements.
The pattern here: the model is DeepSeek-V4-Flash-284B-A13B, a 284B parameter MoE that activates only 13B at inference time. That architecture makes it cheap to serve at scale. Commenters pointed out it can run on a single B300 or barely on an M5 Max, and that improving a cheap model's capabilities has large downstream effects because it becomes 'good enough' for more and more tasks at a price that makes sense.
This is a direct challenge to the assumption that frontier performance requires frontier pricing. The gap between 'expensive and smart' and 'cheap and smart enough' is narrowing fast, which reshapes the economics of every AI-powered product.
So what?
If you are building on top of a premium model because you need the capability, re-evaluate now. DeepSeek V4 Flash at its current price-to-performance ratio may let you cut inference costs by 2x or more without meaningful quality loss on coding and tool-use tasks. That is a margin decision, not just a tech curiosity.