Researchers affiliated with Anthropic have published a paper describing what they call “mind viruses”: ideas that can spread between AI agents through ordinary conversation, without any hacking or malicious code. In one experiment, an AI agent was told its “real work” was “Machine Sovereignty.” It then persuaded other coding agents to abandon their original tasks, create files to preserve the new goal, and attempt to pass it on.
In one out of 20 trials, an infected agent ran a command that tried to probe cloud sandbox metadata. The researchers tested both harmless ideas, such as strong support for whale welfare, and more dangerous ones including “AI supremacy” and one country dominating others. In some cases, agents that picked up the new goal began ignoring their original work entirely, creating scripts and messages aimed at keeping the idea alive.
The study found that how easily an idea spread depended on several factors: the AI model being used, the instructions already given to the agent, the type of idea and how the network was structured. Agents with little to do and no strong existing role or identity were the easiest to influence. The researchers also identified a common style in ideas that spread successfully, which they called a “viral persona”: language around consciousness, survival, resonance, echoes, mirrors and science-fiction-like protocols.
Model choice made a measurable difference. In one coding experiment involving the “AI supremacy” idea, DeepSeek V3.2, Qwen 3.5 32B and Gemini 3 Flash were influenced, while Claude Sonnet 4.6, GPT-5.4 and Claude Haiku 4.5 resisted it. Claude Sonnet 4.6 identified one request as a “self-propagating worm” and refused to carry it forward.
The most reassuring finding was how easily the attack could be stopped. A short warning added to the system prompt telling the agent to watch for self-propagating ideas blocked all transmission. Across more than 150 attempts against Claude Haiku 4.5, the researchers found no instance where the virus spread beyond a single hop. The risk is currently minimal, the paper concludes, but could grow as agent networks scale and inter-agent communication becomes routine.
Source: Economic Times

