LLM Speed Is a Marketing Story, Not a Product Story
Mercury 2.5 hit 770 tokens per second, which would have seemed astonishing a year ago. The HN thread was unimpressed. Commenters pointed out that Cerebras is already at 1,400 tokens per second with a comparable model, that Mercury's benchmark quality sits near the bottom among frontier providers, and that raw speed is meaningless if the outputs are worse. The thread also noted that cheaper open-weight models like DeepSeek and Qwen already undercut Mercury on price for similar capability classes.
The pattern here is that speed as a standalone metric is getting commoditized faster than anyone anticipated. The market is pricing in capability-per-dollar, not tokens-per-second. Inference providers are in a race that increasingly looks like it rewards whoever can offer the best quality at the lowest cost, with speed as a secondary factor only at the extremes.
The contrastive language models thread and the local model replacement thread reinforced this: builders are shopping for the best value across quality, cost, and latency, not optimizing for any single axis.
So what?
If you are building on top of inference APIs, this is a good week to re-evaluate your provider assumptions. The gap between frontier and near-frontier models is narrowing fast on many tasks, and the cost differences are large enough to matter at scale. Do not let benchmark marketing substitute for running your own evals on your actual workload.