The LLM Skeptic Case Is Getting More Articulate
A post titled 'Why I'm still bearish on LLMs after Navier-Stokes' generated real discussion by making a specific, non-denialist argument: even when LLMs solve hard problems like fluid dynamics equations, the best alternative for verifying their output is human review, and human review does not scale to the volume LLMs produce. The post explicitly praises current capabilities while arguing the reliability gap is the fatal flaw.
The same theme surfaces in the Gemini 3.8 thread, where commenters point out that a demo video showed Gemini losing to the most common checkmate pattern in chess. The specific failure mode, confident, fluent wrongness on a verifiable task, is exactly what the Navier-Stokes piece is warning about. These are not isolated dunks; they are evidence accumulating around a coherent critique.
Commenters in the LLM thread are notably tired of both extremes. The denialists who refuse to acknowledge real capability gains and the boosters who hand-wave reliability concerns are both getting less traction. The articulate middle ground, 'impressive but not yet trustworthy at scale without supervision,' is where the serious technical discussion is landing.
So what?
If you are building a product where LLM output goes anywhere near a consequential decision without human review, you are taking on liability your pricing probably does not reflect. The question is not whether LLMs are capable but whether your error-catching infrastructure can keep pace with their output volume.