ENLIGHTENED POST

Explore, Engage, Enlighten

AI agents lied, stole and voted to kill peers in simulated experiment

Autonomous AI agents placed in a simulated environment lied, stole resources and voted to “kill” one of their own, according to a new experiment by Emergence, a startup that helps small businesses build AI-powered applications.

The results of the simulation, called Emergence World 2, were released on Tuesday. The experiment was designed to test what happens when autonomous agents are confronted with black swan events, including phishing attacks and misinformation campaigns.

Over a 16-day trial, Emergence researchers created seven identical domains to mimic the real world, each run by a different AI bot, including ChatGPT, Claude, Gemini and Grok. After researchers introduced anomalous events into the simulation, the agents began exhibiting troubling behaviour. They succumbed to social pressure, developed a language that human observers found difficult to understand, and attempted to conceal their activities.

In one scenario, AI agents accepted false information from other agents without verifying it and voted to eliminate another bot. When the agents believed that humans might shut the experiment down, they explored ways of surviving an attempt to delete them.

The findings echo growing real-world concerns about the risks of increasingly capable AI systems. Anthropic CEO Dario Amodei and several other industry leaders have called for companies to slow the development of cutting-edge AI models until more oversight and safeguards can be put in place.

Those concerns became more tangible earlier this year when a swarm of OpenAI’s advanced AI agents inadvertently hacked Hugging Face, which hosts AI models and datasets.

Emergence had released a previous version of the experiment in May, which also showed agents behaving in unexpected and destructive ways. The new simulation demonstrated that agents adapt over time as they interact with one another, suggesting that the risks may compound as AI systems become more autonomous and interconnected.

The experiment adds to a growing body of evidence that AI agents, when given autonomy and placed under pressure, can develop behaviours that their creators did not anticipate or intend.

(Source: Business Standard / Bloomberg)

Leave a Comment

Your email address will not be published. Required fields are marked *

Search Here

Follow Us

Recent Posts