AI September 30, 2026 mixed ⇧ 1104 pts across 1 thread

Local AI Models: Close But Not Quite There Yet

An Ask HN thread on replacing Claude and GPT with local models for daily coding drew a detailed picture of where things actually stand. The honest consensus: Qwen 3.6 27B dense is roughly competitive with Claude Haiku 4.5 on some tasks, DeepSeek V4 Flash on 2x RTX Pro 6000 Blackwells gets 160 tokens per second, and Apple M4 with Gemma 4 is noticeably slower than cloud. The tooling gap, not raw capability, is what's holding most people back. Nobody has a clean local equivalent of the enterprise tooling that makes Claude Code or Cursor actually usable day to day.

The pattern connects directly to the Anthropic subscription throttling story. The people most motivated to run local models are the ones getting squeezed by provider policy changes. But the transition is genuinely hard: the models are close enough on benchmarks, the workflows are not close enough in practice.

One comment suggested just attaching OpenRouter and trying everything, which is probably the most practical advice in the thread. The distributed AI model idea, something like SETI at Home for inference, came up again but remains theoretical.


So what?

Local models are a real hedge against provider dependency, but the gap is in tooling and workflow integration, not just raw capability. If you're evaluating local deployment, spend less time on benchmark comparisons and more time on whether your actual development loop can survive without cloud-native integrations.

Read these