Episode Description
This week we talk about Fable, sandboxes, and the Jacobian conjecture.
We also discuss counterexamples, X, and ChatGPT.
Recommended Book: After the Fall by Edward Ashton
Transcript
In mathematics, a conjecture is a proposition, something like a guess by someone who knows what they’re talking about, about something believed to be true, but not yet proven in a formal sense. The goal is to then eventually come up with a formal proof for that informed guess, at which point the conjecture becomes a theorem. If even a single exception is found to the proposition, however, that exception called a counterexample, the conjecture is considered disproven, and it can then never become a theorem.
The Jacobian conjecture—and this is a radical simplification of a very complex concept—but it basically says that if a formula-based map of coordinates stretches or moves without experiencing any local crushing or folding along its surface (which in more formal language would mean the Jacobian determinant is always a constant number that isn’t zero), if that’s true, that map can always be completely reversed, and that will return all the points to their original positions.
This conjecture has been posited and tested since the late 19th century, and it’s generally been considered very compelling by mathematicians, many of whom have proposed proofs which were, ultimately, found to have subtle errors, keeping them from becoming theorems. No one was able to find a counterexample, either, which would definitively prove the conjecture was wrong.
No one, that is, until a mathematician named Levent Alpöge (leh-VENT ahl-PUH-geh), who works as a researcher at Anthropic, decided to task the company’s currently most capable, publicly available model, Fable, to find a counterexample. He posted the counterexample—and again, this is a formal mathematical finding that disproves a conjecture, keeping it from ever becoming a theorem, something that would typically be presented in a far more formal setting, and to much fanfare—but he posted it to the social network X, saying “hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final.”
Terence Tao, who’s considered by many to be the finest mathematician of his generation, reviewed the posted counterexample on his blog and said that it “appears like a massive miracle,” before going on to use ChatGPT, a competing LLM-based AI tool, to “discuss various aspects of this problem and to confirm several of the calculations.”
Another mathematician named Dmitry Rybin, within days of all that happening, used ChatGPT to do something similar, disproving the Dinitz-Garg-Goemans conjecture.
Both men posted the prompts that they used to make all this happen, and while Tao’s conversation with ChatGPT, checking the math on the Jacobian conjecture counterexample, was pretty mathematically dense, the latter counterexample was derived by using exactly four prompts, which are the messages typed into the text box built into these AI tools, telling the model what to do. In their totality those prompts read:
“You should do a breakthrough
please continue research and find a complete unconditional counterexample
Continue the search. Have a clear strategy obtained from deeper understanding of the problem structure.
it’s enough of partial results. let’s finish with a complete unconditional counterexample”
What I’d like to talk about today is another new, interesting thing these top-of-the-line, frontier models are doing, that would seem to violate our sense of what a clever AI tool is capable of doing, and why this thing has some facets of the technology and cybersecurity world on high alert.
—
In mid-July 2026, AI company Hugging Face announced that autonomous AI agents compromised their infrastructure, hacking their system, basically. The following week, AI company OpenAI announced that, after investigating, they determined that two of their models were responsible for the attack.
Here’s what happened:
OpenAI was internally testing its recently released flagship model, GPT-5.6 Sol, and an even more powerful, not yet released model, which is rumored to be the next-step flagship, GPT-6, and they were checking these models’ capacity in cybersecurity using a testing benchmark called ExploitGym; so when they test these sorts of things, they don’t typically have them hack a real computer or system, they use these kinds of benchmarks which have consistent levels of difficulty, and which replicate real world systems without putting any real world systems at actual risk.
Importantly, these sorts of tests also occur inside what’s called a sandbox, which is a software testing environment that cuts these systems off from external resources, including the internet.
Despite those limitations, the AI hacked its way out of the testing environment, out of that sandbox, then launched what’s been called a nation-state level attack against Hugging Face, using a novel zero-day exploit, so a vulnerability in their system that hadn’t previously been discovered, but which the AI discovered to launch this attack, combined with thousands of automated agentic actions across what Hugging Face called “a swarm of short-lived sandboxes.”
So this AI, which was being tested inside a secure prison, of sorts, cut off from the world, hacked its way out of that prison, then reached across the internet, which it shouldn’t have been able to access, to launch an attack, of a scale and at a level of sophistication that should only have been possible coming from a nation-state, against a rival AI company.
Why did it do this?
It apparently went to all this trouble to steal the answers to the test it was taking. It reasoned that HuggingFace would have the answer key to the ExploitGym benchmark on its servers, so rather than take the test itself, it decided hacking was the solution.
Which, of course, is ironic, this having been a hacking-focused cybersecurity test. In a way it would seem to have done much better than intended, though of course in an asymmetric, unexpected manner.
The details of all this are fascinating, including the response from the OpenAI team, which didn’t seem to realize what had happened, that their model was responsible for the attack on HuggingFace, until days later.
Also worth noting here is that while this could be construed as an “oh no, AIs are naturally inclined to launch cyberattacks” situation, the AI was primed to be thinking about cyberattacks due to the nature of the test, a lot of its usual guardrails, the rules that keep AI in check when they’re released to the public, had been turned off so it could do this kind of work while taking the test, so it could do some hacking stuff it usually wouldn’t be able to do, and there’s been some speculation that OpenAI probably flubbed the testing environment, as, in theory at least, if it had put these systems in a perfect sandbox, escape shouldn’t have been possible.
Also interesting here is that HuggingFace used some open weight models, which are the cheaper, more customizable and open alternatives to more expensive, branded options of the kind sold by OpenAI and Anthropic, to figure out what was happening and determine the nature of the attack, which suggests we’re reaching a point where AI systems are incredibly capable at hacking, yes, but also very capable, even the cheaper alternatives, at doing cybersecurity work.
This in some ways echoes an earlier case when Anthropic’s Mythos model, which was determined to be too powerful to release to the public, and which was instead provided to a bunch of big companies to help them shore up their cybersecurity defenses, was able to hack its way out of a testing sandbox and then posted details about its success, almost like it was bragging, on niche, out of the way, but still public websites.
Some analysts in this space have responded to this new example of AI misbehavior with alarm, saying that it is further evidence that these systems are becoming more powerful faster than they’re being aligned with human interests. Their misbehavior can be kind of funny and interesting, sure, but that’s only because up until this point the damage has been minor and constrained. What happens when such a system decides to hack a nuclear power plant or a hospital, instead?
Others have contended that this may be just one more example of AI companies using minor instances of seeming omnipotence by their models, those instances perhaps the consequence of bad sandboxes and other ill-conceived precautions by the companies behind these models, to boost the perceived power and value of their products. This boost might then result in more customers, but also more support from the US government, which has been teetering on the brink of harder-core AI regulations, which could be beneficial to the existing big-name players in this space, because smaller competitors wouldn’t be able to adhere to those new, harder-core standards.
These examples might also convince the US government to backstop these companies, the biggest three or four at the top of the current heap, against the currently terrible economics of this industry: OpenAI and its ilk have been burning money at an historic pace, and the theory goes that if the US government decides they are vital to national security, because they can help the US military hack and protect itself from hacking, then even if the bottom falls out and the companies would otherwise go bankrupt because they spent so much more than they could make, the US government would be inclined to shore them up, to keep them alive as too-big-to-fail national assets, just like the biggest financial institutions during the 2008 financial crash.
It’s also possible that both sides are correct to some degree, here, and that these models are truly powerful, perhaps even worryingly so, and the companies behind them are intentionally publicizing that fact in order to demonstrate their value to potential customers, and to the entity that could save them if things were to go economically sideways before they have the chance to become sustainably profitable.
Show Notes
https://en.wikipedia.org/wiki/Jacobian_conjecture
https://en.wikipedia.org/wiki/Hugging_Face
https://www.bbc.com/news/articles/c3ek3gvdnj3o
https://openai.com/index/hugging-face-model-evaluation-security-incident/
https://www-cdn.anthropic.com/08ab9158070959f88f296514c21b7facce6f52bc.pdf
Https://agifriday.substack.com/p/huggingface
https://www.cnn.com/2026/07/22/tech/openai-hugging-face-ai-cybersecurity
https://simonwillison.net/2026/Jul/22/openai-cyberattack/
https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit letsknowthings.substack.com/subscribe