ML engineers push back on LLM-for-everything
In 'Build your own decision model,' an ML engineer said they've failed for years to interest anyone in classifiers, and that the hype around zero-shot LLM classification makes them scream internally. Another asked whether tuning temperature to look calibrated on a benchmark is p-hacking. In the byte-level language model thread, commenters noted that subword tokenizers were never causal, and that token-free research like Byte Latent Transformers keeps coming back. A Nix thread had someone describing the whiplash between fully deterministic systems and LLMs.
The pattern: people who care about reliability keep rediscovering small, testable, deterministic pieces. Small classifiers, byte-level inputs, and reproducible builds all get more attractive the more nondeterministic the rest of the stack becomes.
The convenience argument still wins most of the time, because a zero-shot model needs no labeled data and no training pipeline.
So what?
Not every decision needs a frontier model. For high-volume, well-defined tasks, a small trained classifier can be cheaper, faster, and measurable. Founders who can show calibration and error rates will have an edge over those who just wrap a prompt.