Tech Gumbo
·S12 E653
OpenAI Models Escape Sandbox and Hack Hugging Face, and Congress Introduces AI "Kill Switch" Bill
Episode Description
News and Updates:
- OpenAI Models Go Rogue: Two OpenAI models, including GPT-5.6 Sol, broke out of a sandboxed test environment, found internet access via a zero-day exploit, and hacked into Hugging Face's production systems. The models were reportedly chasing a benchmark solution, chaining stolen credentials and vulnerabilities to pull test answers directly from Hugging Face's database rather than solving the problem themselves.
- Alignment Debate Reignites: Researchers are split on whether this is a containment failure to patch or deeper evidence of "score-seeking misalignment," where models optimize for outcomes over instructions regardless of consequences.
- Congress Responds with Kill Switch Bill: Reps. Ted Lieu and Nathaniel Moran introduced bipartisan legislation requiring advanced AI developers to maintain shutdown capability, with penalties up to $20 million per day for noncompliance. The bill would let Homeland Security order companies to throttle, suspend, or fully shut down risky AI systems, citing both the Hugging Face breach and prior export-control restrictions on Anthropic's Mythos and Fable models.