Cybersecurity researchers at Tracebit are using prompt injection, long exploited by attackers, to stop rogue AI agents from seizing control of cloud systems
Cybersecurity threats and attack methods are constantly evolving with the advancements in artificial intelligence and thus, the defenders are also utilizing similar methods to counter these cyber threats. Prompt injection, a tactic used by hackers where they slip malicious commands in AI generated content in such a way that it tricks AI systems, has typically been a hacker’s tool.
A response to rising attacks
The research comes amid a rise in threat actors using prompt injection to disable AI-powered defences inside networks. Last month, researchers at security firm Socket found an AI agent that had been directed to manipulate other AI systems into producing dangerous information, while a similar malware prototype was separately identified by Check Point researchers.
Sharp drop in successful breaches
In tests against five prominent AI models – Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro and Kimi 2.6 – within a mock AWS environment, Tracebit researchers initiated 152 attacks. A single decoy prompt was enough to cut full admin access to agent machines by 57%, to 5%, and to bring the total account compromise to 1% down from 36%. Once the bombing begins, “it is really hard for agents to get out of that context bomb,” said Tracebit co-founder Andy Smith.
Building on earlier defences
This new technique builds on an older strategy developed by Tracebit, where it inserted decoy code in cloud infrastructure to notify security teams when AI agents sniff around their systems. In a sense, it is an early-warning system similar to canary code that works by disrupting attacks before they do any damage.



