New attack lets hackers plant hidden instructions in AI memory with a single prompt
Brief
A newly demonstrated attack technique allows hackers to plant hidden instructions inside an AI agent’s memory with a single prompt, enabling them to influence how the system responds to future queries. The technique, called InjecMEM, is described in a research paper as a “targeted red-teaming attack paradigm on agent memory systems with just one interaction and no read/edit access to the memory store.”
“The attacker specifies a target topic and target output, aiming to make the agent generate that output for later queries on the topic,” researchers from Shanghai Jiao Tong University and Ant Group wrote in a paper.
The researchers said the method targets AI agents that store past interactions and reuse them during future tasks, enabling malicious content introduced during an initial exchange to persist and affect later outputs.
