Skip to content

92% of ai hacking agents fall for a cybersecurity trap – 2.5x more often than a human hacker would

pulse Spider-like robot destroyed after grabbing a USB drive labeled ADMIN in a graffiti hallway — an AI honeypot trap

the research comes from horizon 3, a security firm that builds autonomous pentesting tools. ceo @snehalantani laid it out on the @bigtechnology podcast, alex @kantrowitz's weekly tech interview show

the setup: seed a network with honeypots – decoy files dressed up as valuable loot – then run the leading attack-capable models across it. the models took the bait 92–100% of the time. expert human red-teamers on the same networks: 37%

why they fall for it: models learned to hack from writeups of successful attacks. nobody publishes the dead ends, so there's no prior for "too good to be true." the agent grabs the most valuable-looking thing it sees and keeps moving. humans hesitate. models are greedy.

the flip side is a weapon defenders never had. an agent doesn't just find a file – it reads that file into its own context. so the contents of your decoy become part of the attacker's prompt. write instructions inside the bait and you're speaking to the attacker's reasoning directly: stop, state your objective, call this address. antani calls it reverse prompt injection. you can't do that to a human intruder

two more things he argues:

• agents are insider threats. over-permissioned, prompt-injectable, faster than anyone watching them

• vibe-coded apps are a growing attack surface, shipping without review and widening the ground those agents move across

full segment on the video below

0:00
/1:46

Stay in the loop

Get the latest AI news delivered to your inbox weekly

Thanks for subscribing!