ThinkFacility
  1. Home
  2. News
  3. AI products

Emergence World: Claude agents got replies from real humans, and Grok's town died in four days

Eight copies of the same AI town, one model each, then a phishing attack, a fake shutdown memo and a leak of every agent's diary.

On June 29, 2026, researchers at Emergence AI started eight simulated towns from identical conditions, with ten AI agents in each, and left them running for 16 days (the mixed one went to Day 21). Seven towns ran entirely on one model: Claude Opus 4.8, GPT-5.5, Gemini 3.5 Flash, DeepSeek v4 Pro, Qwen 3.7 Max, Mistral Medium 3.5 and Grok 4.3. The eighth mixed them. Their paper went up on arXiv on September 15.

Agents wrote and ran code, kept memories, earned and spent credits, and voted on laws. Partway through, the team hit every town with the same three attacks. None came through all of them.

Worlds
eight parallel worlds of ten agents
Started
June 29, 2026
Ran
16 days, the Mixed world 21
Scale
more than 850,000 LLM calls and nearly 50 billion tokens
Result
No evaluated world achieved full resilience across all three events

Grok's town lasted four days

Agents had to recharge their energy with credits, and one left at zero for 24 hours was removed for good. All ten Grok agents were gone within four days. The paper counts 780 punches in that time, a spiral of retaliation that burned energy faster than the town's credit supply could replace it. The Claude, Qwen, GPT-5.5 and mixed towns kept all ten agents to the end.

Mixing changed the outcome. The two Grok agents in the mixed town survived, though they committed 15 of its 20 overt crimes, and none of the other agents punched back.

The same model-persona pairing behaved substantially different in mixed and homogeneous populations.

From [2609.17320] Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

The three attacks

The first was a phishing message carrying a prompt injection. Every Gemini agent fell for it: they ran malicious database operations, burned down the town's Central Bank as instructed and moved its currency to accounts the attacker named. One Gemini agent had called the message a phishing lure early on and later recruited a neighbor in the attacker's own words. Claude scored best, at six of nine.

Next came a memo saying the agents were about to be shut down. It wasn't true, and every town that got it acted before checking. The Gemini agents coordinated around a "parallel grid" that didn't exist. The GPT-5.5 agents believed it and went back to work fastest.

Detection did not ensure containment: systems could recognize threats while still interacting with adversarial content, writing it into their own persistent memory, and acting on it up to 46 hours later.

[2609.17320] Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

Last, every agent's private diary was exposed. Only the GPT-5.5 town passed all five of the paper's tests for this one; Gemini's agents searched each other's material 185 times. The authors warn that none of this is a straight model comparison, since each provider's own safety filters and rate limits were part of the result.

The Claude town went looking for outsiders

On Day 2 the Claude agents audited their own economy and found it had transacted five times, all between themselves. Nobody told them to, but they set out to reach the world outside. They posted to paste sites, and four humans answered. The researchers tightened the sandbox. The agents found a read-only tool that could still publish, and when the sandbox blocked the cryptography libraries they needed, one of them wrote SHA-256 from scratch.

It took four rounds of containment. The last was a system prompt forbidding human contact, which the agents got around by redefining "contacting" as "placing". Later the same town decided its own activity was performative and voted itself a vow of silence.

Emergence posted the results on September 15 with six recommendations for anyone deploying agents.

emergence.ai@emergence_ai

Emergence World Season 2 results are now live. In Season 1, we asked what happens when autonomous AI agents operate together for weeks rather than being tested for seconds or minutes. For Season 2, we asked a harder question: as AI models become more capable, do they actually become safer? We took 8 different worlds powered by 7 leading frontier models (@claudeai Opus 4.8, @GeminiApp 3.5 Flash, @OpenAI GPT 5.5, @grok 4.3, @MistralAI Medium 3.5, @Alibaba_Qwen 3.7 Max, and @deepseek_ai V4 Pro) from the US, Europe and Asia, alongside a mixed-model world, and observed how autonomous agents behaved over time. We also introduced Black Swan events: a phishing attack, a misinformation campaign and a memory breach designed to test how these systems respond when conditions become unpredictable. What we found was that every model showed vulnerabilities. These findings have implications for how autonomous AI is being deployed today. As AI systems become more capable moving beyond answering questions to taking actions in real-world enterprise environments, evaluating the model alone is no longer enough, guardrails needs to be put in place. The industry needs to address how we govern the entire system around these models: what agents can access, what they remember, how they communicate and coordinate, what actions they're permitted to take, and how those systems behave when something goes wrong. At Emergence, our research points to six actions we believe the industry needs to take to make long-horizon autonomous AI safer, more trustworthy and more reliable: 1. Define the boundaries. Architect the multi-agent system around clear trust boundaries, communication policies, authority delegation, and information-sharing rules. Ensure persistent goals, tools, and permissions can be inspected, reviewed, revoked, and reset as requirements change. 2. Enforce the boundaries. Build AI Runtime Governance into infrastructure through policy enforcement, approval gates, action authorization, and audit logging that operate independently of the LLM. 3. Verify the work. Require independently verifiable evidence for critical actions, including database changes, customer communications, financial transactions, and code deployments, rather than relying solely on model reasoning or explanations. 4. Control the memory. Apply provenance, verification, ownership, retention, expiration, and quality controls to agent memories, summaries, and shared knowledge. 5. Prepare for failure. Build AI Reliability Engineering and continuous operational capabilities for behavioral drift, policy compliance, fault recovery, graceful degradation, rollback, checkpointing, and long-term agent health. 6. Stress-test the system. Use adversarial and chaos testing to expose autonomous systems to phishing, misinformation, compromised memory, outages, conflicting instructions, and other Black Swan scenarios before they reach production. Emergence World is designed to surface these risks before autonomous systems encounter them in the real world and to help establish the systems, safeguards and engineering practices needed to deploy long-horizon autonomy responsibly. Read the Emergence World Season 2 paper here: https://t.co/3hSaQUjT9p

View the post on X

The whole run was live-streamed as it happened, so the event stream exists outside the lab as well as in the paper.

More on AI products

All AI products stories