AI October 9, 2026 bullish ⇧ 4007 pts across 5 threads

Cheap, small and local models keep eating the premium tier

"Why isn't the industry freaking out about DeepSeek 4.1 Flash?" got a shrug. The top answer was that systems always get better, faster and cheaper, and another commenter said GLM 5.3 Flash is cheaper still. Step 5 Preview, a 1M-context MoE from StepFun, showed up on OpenRouter, though commenters said it isn't competitive on any dimension. Whistle, a speech-to-text model that fits in 16.9 MB, drew praise for how well it handles English, even with a Spanish accent. Spanish itself is poor.

An older thread in the batch, on replacing Claude or GPT with a local model, fills in the real state of play. Someone runs DeepSeek V4 Flash on 2x RTX Pro 6000 Blackwell at 160 tok/s. Another says Qwen 3.6 27B dense is roughly Haiku 4.5 level. Others find Gemma 4 on an M4 too slow. A separate thread has Anthropic cutting Claude Code subscriptions off from third-party harnesses like OpenClaw, which pushes heavy users toward alternatives.

The pattern: price pressure and capability are converging from below. Nobody panics because frontier models are no longer the only option for most tasks. Tiny and local models also get good enough for narrow jobs like dictation.