Cheap Fine-Tuning Is Eating Frontier Model Revenue
Ramp spent $500 fine-tuning a 9B open model for catalog review and beat GPT-4-class frontier models on their specific task. The thread around this story is active and pointed, with people noting the obvious implication: for narrow, well-defined tasks, the economics of fine-tuning a small open model are so favorable that paying per-token for a frontier model becomes hard to justify.
The pattern here is not just cost. It is that the fine-tuning workflow itself is getting easier and cheaper at the same time that open models are getting stronger. A few months ago this comparison would have been embarrassing for the open model. Now it is embarrassing for the frontier labs. The benchmark thread on Opus 5 reinforces this, with commenters saying they felt no 'wow factor' compared to Opus 4, suggesting diminishing returns at the top.
The counterpoint raised in the thread is real: companies with 2x revenue have money to spend on AI experimentation, and correlation with business outcomes is not causation. But even if you discount the Ramp story, the directional trend is clear.
So what?
If you are building a product that calls a frontier API for a task you do repeatedly at scale, you should be running this calculation now. Fine-tuning a smaller open model for your specific domain is increasingly likely to win on quality and will almost certainly win on cost. The window where frontier models are the obvious default choice for everything is closing.