Open weights and tiny models keep closing the gap
Mistral Large 4 landed as a 1T-parameter open-weight model, with claims of parity with GLM 5.3 on DeepSWE. One commenter said Mistral had "slightly proved me wrong" and wasn't mad about it. Google released EmbeddingGemma 2, a multimodal embedding model under Apache 2.0, and commenters said the license is the point. Their argument: you shouldn't build on a closed, hosted-only embedding model. Strands Decider 2B is a small open decision model, and one user already runs it on-device in a Chrome extension to filter email.
The local-coding thread shows how close this is to practical. One user reports Qwen 3.6 27B dense performing around Claude Haiku 4.5, maybe Sonnet on some tasks. Another runs DeepSeek V4 Flash on two RTX Pro 6000 cards at 160 tok/s. The pattern: the frontier is no longer the only interesting place. Specialized small models and permissive licenses are where builders are actually shipping.
The caveat is speed and tooling. Gemma 4 on an M4 was too slow for one person, and others point to the lack of enterprise tooling for choosing and running local models. The benchmark claims for Mistral also still need independent confirmation.
So what?
If your product depends on one closed API for classification, routing, or embeddings, you now have a credible open alternative that costs less and can't be revoked. Test a small open model for those narrow jobs this month. Keep the frontier model for the hard reasoning only.
Read these
Mistral Large 4
EmbeddingGemma 2: An open, lightweight multimodal embedding model
Strands Decider 2B: a small, open-source, decision model
Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?