Model pricing is the battleground, and local models are creeping up
Gemini 4 Argon launches at $2 per million input tokens and $10 per million output, with cached input at 95% off. The thread's reaction was muted: Gemini isn't beating the leaders on quality, and one commenter called it the model that is "routinely borderline psychotic." Google is still holding back developer, enterprise and consumer availability while it tunes guardrails. So the pitch is price and caching, not a capability jump.
The same pressure shows up from the other direction. In the thread on Anthropic ending Claude Code subscription use with OpenClaw, users quoted the "outsized strain" line back at the company: if you pay for tokens, why is using them a problem? One user said they hit rate limits within 1:45 on afternoons and want a tier that doesn't lock them out. Meanwhile the Ask HN on replacing Claude or GPT with a local model has real answers. One person says Qwen 3.6 27b dense is about equal to Haiku 4.5. Another runs DeepSeek V4 Flash at 160 tok/s on two RTX Pro 6000 Blackwell cards. Magnitude (YC S25) launched a self-optimizing inference engine for agents, and the first question was the business model.
The pattern here: subscription flat-rate pricing is breaking under agent workloads, labs are fencing off third-party harnesses, and the escape hatch is open-weight models that are good enough for routine coding. Gemma 4 on an M4 is still too slow for people, so local isn't there for everyone yet.
So what?
Don't build your margins on one lab's subscription terms. Design your product so you can swap models and route cheap work to cached or open-weight options. Cached-token pricing like Argon's 95% discount rewards products that reuse long prompts, so structure yours that way.
Read these
Gemini 4 Argon
Tell HN: Anthropic no longer allowing Claude Code subscriptions to use OpenClaw
Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents