Claude Opus 5's Verified Benchmark Win, and Why 'Half Price' Is a Trap

July 25
21 mins

View Transcript

Episode Description

Anthropic's new model posted a rare thing today: a benchmark win an outside group independently verified -- a roughly four-fold jump on the hardest test of figuring out an unfamiliar world from scratch. But the 'half price' story is a trap, and the number hides more than it shows. We take three papers from the same day and use them to crack open what a benchmark actually measures: a top model that aces leaderboards yet can't count patches of grass in a photo, a tiny research agent that beats models many times its size by remembering its own dead ends, and a data trick that's real but wildly oversold. One thread runs through all three -- a score is a claim about one narrow skill under one setup, and what matters is what a model does with what it's got.

Chapters

0:00 Cold Open -- The Benchmark Is Real, The Price Cut Isn't
1:29 The Headlines
5:53 Intro
7:14 The Test Where Thinking Harder Doesn't Help
11:38 The Small Agent That Wins By Remembering What It Got Wrong
15:41 Why A Model Fed Almost No Data Can Punch Above Its Weight
19:54 Wrap-Up

Links

Cold Open -- The Benchmark Is Real, The Price Cut Isn't -- https://www.anthropic.com/news/claude-opus-5
The Test Where Thinking Harder Doesn't Help -- https://arxiv.org/abs/2607.16165
The Small Agent That Wins By Remembering What It Got Wrong -- https://arxiv.org/abs/2607.21461
Why A Model Fed Almost No Data Can Punch Above Its Weight -- https://arxiv.org/abs/2510.04071
See all episodes