AI Release Flood Makes Model Selection Intractable
In a single day: Google shipped Gemini 3.7 Flash at $0.75/1M input tokens, OpenAI announced GPT-5.6 Sol Ultrafast, Chinese lab Zhipu dropped GLM-5.3 with claimed frontier coding scores, Mistral released OCR 4.1, and DeepSeek updated its peak/off-peak pricing. HN commenters on the GLM thread said it plainly: 'A flood of releases today, really difficult to make out for someone who does not use or test all these models on complex real world use cases as to how people decide which ones to use (besides price).'
The pattern here is that model releases are now happening faster than the evaluation infrastructure can keep up with. Benchmarks like the Cognition FrontierCode leaderboard are getting cited in arguments between labs, but commenters flag that the same benchmark puts labs in very different positions than each other's marketing implies. Mistral is getting called out for being expensive relative to open alternatives, and one commenter bluntly said 'The Chinese did it better, Mistral is alive thanks to regulations.'
DeepSeek's pricing update adds another wrinkle: peak-hour pricing jumped from $0.28/M to $1.32/M output tokens, and commenters noticed that peak hours are Chinese business hours, suggesting DeepSeek's actual customer base is still mostly domestic. That changes the calculus for Western builders who assumed DeepSeek was primarily targeting them.
So what?
Founders building on top of any single model are now exposed to price and capability volatility on a weekly cadence. The practical move is abstracting your model calls behind a routing layer so you can swap providers without code changes. Betting on any one lab's pricing as stable for more than a quarter is now a bad assumption.
Read these
Gemini 3.7 Flash
Accelerating GPT-5.6 Sol Ultrafast
GLM-5.3: Frontier coding with emergent cyber capabilities
Mistral OCR 4.1
DeepSeek peak/off-peak pricing update