Researchers Warn: New Attack Can Manipulate AI's Memory
Researchers have warned of a new security vulnerability that could pose one of the most serious challenges facing modern AI systems, after they successfully developed an attack named GhostWriter, capable of planting fake memories within the memory of AI assistants, potentially leading them to make incorrect decisions or execute malicious commands long after the attack occurs.
This warning comes at a time when AI companies are racing to develop digital assistants with long-term memory, allowing them to remember users' preferences, habits, and projects to provide a more personalized experience.
The attack does not hack the system... but changes its memory
The study was conducted by researchers from New Mexico State University, who explained that the GhostWriter attack does not rely on hacking the AI model or directly stealing data, but rather targets manipulating what the system remembers, according to a report published by Digital Trends and reviewed by Al Arabiya Business.
Instead of attacking the model itself, the attacker injects misleading information into the assistant's long-term memory through hidden instructions or untrusted external content.
This fake information remains stored in the memory until the system retrieves it later while executing a seemingly normal request, influencing its decisions without the user or the system itself being aware of the manipulation.
How does GhostWriter work?
According to the researchers, the attack goes through two main stages:
Memory injection: Malicious information is inserted into the assistant's memory in a covert manner.
Attack activation: The system later retrieves that information when responding to a real user request, acting based on fake data it believes to be true.
For example, if a user asks their AI assistant to summarize emails from the bank, a contaminated memory could cause the assistant to secretly send those emails to an attacker, use incorrect contact information, remember wrong deadlines, or even adopt fake preferences previously planted in its memory.
More dangerous than 'prompt injection' attacks
The researchers emphasize that GhostWriter differs from traditional prompt injection attacks, whose effects are usually limited to the current conversation.
In this scenario, the malicious information remains stored in long-term memory and continues to affect the assistant's behavior across multiple sessions until it is detected and removed.
AI memory: the new attack surface
This study comes at a time when memory has become one of the most important features that AI companies are competing over, as most major companies are developing assistants capable of remembering users for weeks, months, or even years.
This gives AI assistants greater ability to understand context and provide more accurate responses, but at the same time opens a new door for cyberattacks.
During tests, the GhostWriter attack achieved a success rate of about 98% in planting information into memory, while the fake memories succeeded in influencing the decisions of the latest AI agents by approximately 60% when later retrieved.
The researchers believe these results indicate that current memory management systems are still unable to reliably distinguish between correct information and manipulated data.
New defensive framework
The research team did not stop at revealing the vulnerability; they also proposed a new defensive system called Agentic Memory Sentry (AM-Sentry).
This system relies on scanning information before storing it in memory, along with implementing stricter policies for managing stored memories.
According to the researchers, this framework significantly reduced the chances of a successful GhostWriter attack, while maintaining the efficiency of AI assistants and their ability to perform tasks.
The next security battle
The researchers conclude their study by affirming that the future of AI security will not be limited to protecting models from malicious commands, but will also extend to protecting their memory from manipulation and rewriting.
As AI assistants evolve to handle tasks such as email management, meeting scheduling, code writing, and making decisions on behalf of users, protecting what AI 'remembers' may become as important as protecting the data it produces.
Ad material
Ad material
Original source: Al Arabiya
Comments (0)
Be the first to comment.