AI October 5, 2026 bullish ⇧ 3045 pts across 3 threads

Local models are closing in on cloud coding

A user posted that Qwen 3.8 Flash Next (125B) runs at about 124 tokens per second on a 4090 with 128GB of DDR5 and a Ryzen 7950x3d, using a mixture-of-experts offload trick (49953495). The first replies went straight to the weak spots. One asked for benchmarks on the quantized versions. Another pointed out that generation speed is the easy half, and that prompt processing at 16k context is the real test.

The Ask HN on whether anyone has replaced Claude or GPT with a local model for daily coding (48542100) gave mixed answers. One person says Qwen 3.6 27B dense is about equal to Haiku 4.5. Another runs DeepSeek V4 Flash on two RTX Pro 6000 Blackwell cards at 160 tok/s. A third tried Gemma 4 on an M4 and found it too slow. The Anthropic thread (47633396) is the push behind all this. Anthropic stopped letting Claude Code subscriptions power OpenClaw, calling the usage an 'outsized strain,' and users are angry about rate limits that hit them within 1:45 on afternoons.

The pattern here: every time a vendor tightens limits, the local option gets more attractive. The gap is no longer 'can it work at all.' It is now speed, prompt processing, and tooling.


So what?

If your product depends on one lab's subscription or API terms, you carry platform risk. Build your agent layer so the model can be swapped, and test it against an open-weight model now. Local inference is also a real selling point for customers who can't send code or data to the cloud.

Read these