Your Model Fits -- That Was Never The Hard Part: Inside the New Local-AI Wall

August 3
23 mins

View Transcript

Episode Description

Someone ran a 284-billion-parameter model in about five gigabytes this week, and half the internet learned the wrong lesson. Fitting the weights stopped being the constraint -- and the moment it did, three quieter problems walked in: storing a model's memory in shorthand changes which words it picks, the same speed switch makes one machine faster and another slower, and making a video model fast also makes it boring. We open up DeepSeek V4 Flash's cache contract, the DSpark drafter now sealed inside the checkpoint, and the reverse-KL reason fast video gets prettier and samey at once -- every mechanism traced to the papers.

Chapters

0:00 The wrong lesson from a five-gigabyte giant
1:17 The headlines
6:31 Intro
7:51 Why shorthand changed its mind
12:00 The drafter shipped inside the box
17:08 Prettier and more boring at once
22:10 Wrap-up

Links

The wrong lesson from a five-gigabyte giant -- https://github.com/NeelM0906/Mference
Why shorthand changed its mind -- https://arxiv.org/abs/2606.19348
The drafter shipped inside the box -- https://arxiv.org/abs/2607.05147
Prettier and more boring at once -- https://arxiv.org/abs/2607.26811
See all episodes