Tiny local decision models as a new infra primitive
Two related projects landed on HN today: Jeff, a set of small Qwen3.5 and Gemma fine-tunes for zero-shot classification running at ~30ms locally, and Jeeves, which adds reasoning to Jev-like decision models. Both are explicitly designed to be dropped into code as a decision layer rather than used as general-purpose assistants. Jeff is positioned as a locally runnable, fine-tunable alternative to cloud classification APIs.
The pattern here is the emergence of sub-1B parameter models as a serious infrastructure primitive for classification and routing tasks. These are not trying to be ChatGPT. They are trying to be fast, cheap, local, and correct on a narrow task. The HN comments show real tension between the appeal of the concept and the execution: one commenter reported 70% accuracy vs 94% for a competing approach on their specific task, which for classification workloads is a deal-breaker.
Separately, the ESP32S3 cluster running a 1.58-bit BitNet language model is the hobbyist extreme of the same trend: ML inference on genuinely constrained hardware. The joke comment about 'AI in every lightbulb running Kubernetes' lands because the underlying trajectory is real.
So what?
For founders building products that need fast yes/no or classification decisions, locally runnable sub-1B models are approaching the point where they are worth evaluating against cloud APIs. The cost and latency advantages are real. The accuracy ceiling is the current limiting factor. Run your own benchmarks on your specific task before committing to either direction.