AI Agent APIs Arrive, Trust Questions Follow Immediately
OpenAI shipped an Agents API and Cognition launched SWE-2, a new coding-focused model post-trained from Kimi K3 that claims to rival Fable 5.1 and GPT-Astra on certain benchmarks. The threads around both are instructive. The OpenAI Agents API thread's top comment is 'Perfect for when you want your data to be stolen programmatically,' which is a jokey but pointed summary of the trust problem. The SWE-2 thread notes that Cognition is 'flying under the radar' while Anthropic and OpenAI dominate headlines.
The benchmark skepticism on SWE-2 is real and worth tracking. Commenters immediately flag that the benchmark selection looks curated, Terminal-Bench 4 results are weak, and other standard benchmarks are absent. This is a recurring pattern in the AI coding agent space: every new model launches with cherry-picked benchmarks, and the HN community has gotten good at spotting the gaps.
Taken together, the two stories reflect the current state of the agent market: infrastructure is shipping fast, trust and evaluation rigor are lagging, and the gap between benchmark claims and real-world performance remains wide.
So what?
If you're evaluating AI coding agents for your team, benchmark press releases are nearly useless. Build your own eval on tasks representative of your actual codebase before committing. The gap between marketing benchmarks and production performance is still large enough to matter.