OpenAI's Model Optimized Its Own Serving Stack -- and the Quiet Efficiency Race Underneath It

July 30
18 mins

View Transcript

Episode Description

OpenAI says its newest model rewrote the low-level code that runs it and cut serving costs by about a fifth -- and the internet called it the singularity. It isn't, and the real story is bigger: three research teams shipped the same quiet idea this week -- stop paying twice for work you already did. We get into a frozen 12B model that answers already-solved problems at zero cost, an NVIDIA kernel that makes video generation roughly twice as fast for free, and a training trick that catches a model's wrong turn before it wastes three pages. Plus the FCC putting every foreign-made robot on a national-security list, and Meta's data-center build swallowing 98% of its cash flow.

Chapters

0:00 The AI That Rewrote Its Own Plumbing
1:51 The Headlines
6:41 Intro
7:43 Zero Tokens, Forever
11:13 Skipping The Boring Parts
14:25 Catch The Wrong Turn Early
17:34 Stop Paying Twice

Links

The AI That Rewrote Its Own Plumbing -- https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/
Zero Tokens, Forever -- https://arxiv.org/abs/2607.23806
Skipping The Boring Parts -- https://arxiv.org/abs/2607.24027
Catch The Wrong Turn Early -- https://arxiv.org/abs/2607.26057
See all episodes