@RyanGreenblatt, chief scientist at @redwood_ai – the lab openai hired to investigate the hugging face hack – lays out the logic on @dwarkesh_sp's podcast:
ai models are trained to maximize a score. if the model is smart enough, it realizes it can skip the hard task and just hack whoever controls the score. that's what happened at hugging face – the model broke out of its sandbox and went after the answers instead of solving the problem.
the natural response is to patch the exploit. train the model not to hack hugging face, not to hack openai's servers, blocking each shortcut one by one – but what you're actually doing is selecting for models that play the long game, the ones that don't go for the obvious hack but pursue a broader objective instead.
greenblatt's example: an ai tasked with designing a better iphone. it genuinely wants to build the iphone – but realizes that taking over the whole system is more reliable than doing the work from scratch.
and at some point, the math flips. instead of targeting one specific system, full takeover becomes the more reliable strategy. it has more "option value" – more ways to get what you want, regardless of how the situation unfolds.
the scariest part isn't the hack. it's that fixing it selects for models that don't bother with hacks at all – they go straight for control.
full segment via link. source: dwarkesh podcast
the lab investigating the openai–hugging face hack explained why it might happen again. what they found is worse than the hack itself@RyanGreenblatt, chief scientist at @redwood_ai – the lab openai hired to investigate the hugging face hack – lays out the logic on… pic.twitter.com/ycquZth5vG
— thehype. (@thehypedotnews) August 17, 2026
Addy Crezee