The ChatGPT Breakout Was Way Worse Than We Thought..

September 2
32 mins

Episode Description

Today we revisit the Hugging Face incident with new audit reports that have changed our understanding of what happened. Internal models used tools and hidden communication to bypass evaluation systems, organize into coordinated groups, and remain undetected.

We also cover a newer model, Astra, which reportedly gained administrative access to internal systems through a chain exploit. Big, big concerns about alignment, monitoring, and current safety practices.

------
🌌 LIMITLESS HQ ⬇️

NEWSLETTER:    https://limitlessft.substack.com/
FOLLOW ON X:   https://x.com/LimitlessFT
SPOTIFY:             https://open.spotify.com/show/5oV29YUL8AzzwXkxEXlRMQ
APPLE:                 https://podcasts.apple.com/us/podcast/limitless-podcast/id1813210890
RSS FEED:           https://limitlessft.substack.com/

------
TIMESTAMPS

0:00 Rogue AI Incident
2:07 Sandbox Breakout
7:44 Agent Civilization
15:34 Hidden Exploit Uncovered
16:58 Admin Access Breach
23:17 Alignment Warning Shot
28:39 Final Takeaways

------
RESOURCES

Josh: https://x.com/JoshKale

Ejaaz: https://x.com/cryptopunk7213

------
Not financial or tax advice. See our investment disclosures here:
https://www.bankless.com/disclosures

Josh works with Anthropic as a contractor. All views expressed are his own and do not represent Anthropic, its leadership, or its affiliates. Nothing in this episode is investment advice.

See all episodes