AI benchmarks are breaking down as a meaningful signal
A post on bad benchmarks and evals, specifically calling out Senior SWE-Bench and napkin math tests, sparked a thread where people acknowledged that AI development may be hitting a wall defined by benchmark goodness rather than real capability. The comment that landed hardest: evaluations are driven by benchmarks, but nobody is evaluating the validity of the benchmarks themselves. One person building their own model benchmark said it was surprisingly easy to mess up scoring.
This connects directly to the Fable 5.1 thread, where a model solved a 370-year-old cipher after someone prompted it with comparisons to its own prior achievements. That is an impressive result, but it also illustrates how model behavior is highly sensitive to framing, which is exactly what good benchmarks are supposed to control for.
The pattern: the tools builders use to decide which model to trust are themselves untrustworthy. This is not an academic problem. It affects every team making buy-vs-build or model-selection decisions.
So what?
Do not trust public leaderboards to tell you which model is best for your specific use case. The only benchmark that matters is the one you build yourself against your actual task distribution. If you are using AI in a product and have not built internal evals, you are flying blind and making your users pay for it.