Can an AI Actually Hack? What Two Safety Institutes Finally Measured

July 24
22 mins

View Transcript

Episode Description

Everyone has an opinion on whether AI can hack -- almost nobody had a number, until two government safety institutes went and measured it. We climb the 16-rung ladder that separates crashing a program from owning it, why the grading can't be faked, and the unsettling twist: today's public models stall at a wall that one private model already walked through. Then the sentence that detonates the whole comfort -- an open model being 'below the frontier' is a statement about budget, not brains, and budget is exactly the leash you drop when you release the weights. We close on a training paper whose author honestly corrects his own headline. Full daily rundown at groundtruth.day.

Chapters

0:00 The Kill Switch That Might Not Cover Its Own Example
1:14 The Headlines
6:27 Intro -- Measuring the Scary Part
7:44 Can An AI Actually Hack? How You'd Even Measure It
12:55 Below The Frontier, Or Just Under Budget?
16:51 Which Parameters Deserve To Remember?
21:01 Wrap-Up

Links

The Kill Switch That Might Not Cover Its Own Example -- https://lieu.house.gov/media-center/press-releases/reps-lieu-and-moran-introduce-bill-require-kill-switch-ai-systems-can
Can An AI Actually Hack? How You'd Even Measure It -- https://arxiv.org/abs/2605.14153
Below The Frontier, Or Just Under Budget? -- https://arxiv.org/abs/2603.11214
Which Parameters Deserve To Remember? -- https://arxiv.org/abs/2607.19058
See all episodes