We've Been Using GPT-6 Astra for a Few Weeks, Here's What We Think...

September 3
1h 51m

View Transcript

Episode Description

Theo & Ben break down OpenAI's latest model, Astra, and why it's their new benchmark for AI coding, computer use, and multimodal work. But it's not all good: they explain why its UI, stopping behavior, and agent reliability still create friction in real software workflows. From DEFCON puzzles to coding-agent PRs, we're comparing Astra with Fable and asking what a trustworthy OpenAI model should do next.


Thank you to PostHog for sponsoring today's episode!

Listen wherever you get your podcasts:

Sources available on our Substack:

Timestamps

00:00 Meet Astra

04:03 Best Model Ever, With Catches

06:33 Reasoning and 3D Benchmarks

20:00 Astra Rebuilds Ping.gg

30:04 Instruction-Following Problems

44:47 The Uncommitted Fix Debate

01:09:37 The PR Babysitting Failure

01:22:41 Why Fable Still Wins

01:30:41 Multimodal and Computer Use

01:39:29 Astra vs. Fable Fleet Data

See all episodes