Local models are close, but not close enough
Several threads circled the same question: can you drop the paid API and run your own model? An Ask HN on replacing Claude or GPT with a local model for daily coding got honest mixed answers. Gemma 4 on an Apple M4 was too slow. One user runs DeepSeek V4 Flash on two RTX Pro 6000 Blackwell cards at 160 tokens per second. Another says Qwen 3.6 27B dense is about Haiku 4.5 level, maybe Sonnet on some tasks. Janus, a Go binary that runs GGUF models through Vulkan on AMD, Intel and Nvidia, showed up as a Show HN, and people immediately asked for benchmarks against vLLM and SGLang. Someone also noted that Vulkan adds a lot of overhead on Intel hardware. On the Pi 1.0 thread, a commenter said they will stay on opencode until a local model is good enough for coding.
Behind this sits the Anthropic story. Anthropic stopped Claude Code subscriptions from working with OpenClaw, citing "outsized strain", and users were hitting rate limits within hours. When a vendor can change the terms on your workflow overnight, running your own model starts to look like insurance.
The counterpoint: the people doing it successfully need very expensive hardware, or accept a model one tier below the frontier.
So what?
Local inference is not a replacement yet, but it is a credible fallback and a negotiating position. Build your tooling so you can swap in an open-weight model, because vendor pricing and access rules keep shifting under you.
Read these
Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia
Tell HN: Anthropic no longer allowing Claude Code subscriptions to use OpenClaw
Pi 1.0