the research comes from horizon 3, a security firm that builds autonomous pentesting tools. ceo @snehalantani laid it out on the @bigtechnology podcast, alex @kantrowitz's weekly tech interview show
the setup: seed a network with honeypots – decoy files dressed up as valuable loot – then run the leading attack-capable models across it. the models took the bait 92–100% of the time. expert human red-teamers on the same networks: 37%
why they fall for it: models learned to hack from writeups of successful attacks. nobody publishes the dead ends, so there's no prior for "too good to be true." the agent grabs the most valuable-looking thing it sees and keeps moving. humans hesitate. models are greedy.
the flip side is a weapon defenders never had. an agent doesn't just find a file – it reads that file into its own context. so the contents of your decoy become part of the attacker's prompt. write instructions inside the bait and you're speaking to the attacker's reasoning directly: stop, state your objective, call this address. antani calls it reverse prompt injection. you can't do that to a human intruder
two more things he argues:
• agents are insider threats. over-permissioned, prompt-injectable, faster than anyone watching them
• vibe-coded apps are a growing attack surface, shipping without review and widening the ground those agents move across
full segment on the video below
92% of ai hacking agents fall for a cybersecurity trap – 2.5x more often than a human hacker would
0:00
/1:46
Addy Crezee