Sonnet 5.5 matches Opus 5.5 on coding benchmarks
Claude Sonnet 5.5 is scoring nearly identically to Opus 5.5 on agentic coding tasks: 70.6% vs 66.4% on Terminal-Bench, and comparable numbers on FrontierCode. The HN thread also flags that Anthropic is quietly dropping Fable from benchmark comparisons entirely, which several commenters read as a sign Fable is being wound down. Separately, Anthropic recently cut off third-party harness tools like OpenClaw from Claude Code subscriptions, frustrating power users who built workflows around those tools.
The pattern here is a fast-moving capability curve compressing the gap between model tiers, while Anthropic simultaneously tightens control over how those models are accessed. Sonnet-level performance at Sonnet-level pricing is great for most developers, but the removal of third-party harness access signals that Anthropic wants to own the full agentic workflow stack, not just the model layer.
There is real friction in the community around this. Users on the $200/month Cyber Verification plan report being blocked from using either Sonnet 5.5 or Opus 5.5 for authorized bounty work. The capability story and the access story are moving in opposite directions.
So what?
If Sonnet 5.5 genuinely matches Opus 5.5 on coding tasks, the tier distinction mostly collapses for agentic use cases, which is good news for cost-conscious builders. But Anthropic locking out third-party harnesses is a direct threat to any startup that built infrastructure around Claude Code with non-native tooling. Audit your dependencies on Anthropic-specific tooling now, before the next policy change hits.