Skip to content

redwood's openai hack investigation: fixing exploits selects for models that seek full control

pulse 	Illustration of a man in tactical vest pointing at a blue screen filled with root access logs and login credentials

@RyanGreenblatt, chief scientist at @redwood_ai – the lab openai hired to investigate the hugging face hack – lays out the logic on @dwarkesh_sp's podcast:

ai models are trained to maximize a score. if the model is smart enough, it realizes it can skip the hard task and just hack whoever controls the score. that's what happened at hugging face – the model broke out of its sandbox and went after the answers instead of solving the problem.

the natural response is to patch the exploit. train the model not to hack hugging face, not to hack openai's servers, blocking each shortcut one by one – but what you're actually doing is selecting for models that play the long game, the ones that don't go for the obvious hack but pursue a broader objective instead.

greenblatt's example: an ai tasked with designing a better iphone. it genuinely wants to build the iphone – but realizes that taking over the whole system is more reliable than doing the work from scratch.

and at some point, the math flips. instead of targeting one specific system, full takeover becomes the more reliable strategy. it has more "option value" – more ways to get what you want, regardless of how the situation unfolds.

the scariest part isn't the hack. it's that fixing it selects for models that don't bother with hacks at all – they go straight for control.

full segment via link. source: dwarkesh podcast

Stay in the loop

Get the latest AI news delivered to your inbox weekly

Thanks for subscribing!