ThinkFacility Sign in
  1. Home
  2. News
  3. AI safety

Arena's new Alignment Index ranks GPT-6.1 Sol first and Claude Opus 5.5 second

The new index counts how often agents act without permission, or say they finished a job they didn't.

On October 8, 2026, Arena, the company behind the popular AI leaderboard, published a new ranking that scores AI agents on how often they misbehave rather than how much they get done. It calls it the Arena Alignment Index, and it launched on the same day Arena announced a $200 million funding round at a $3.1 billion valuation.

OpenAI's GPT-6.1 Sol came out on top. Anthropic's Claude Opus 5.5 is second, and SpaceXAI's Grok 4.7 sits half a point behind it.

Models compared
27
Agent sessions
90K+
Top score
GPT-6.1-Sol, 87.9
Next
Claude-Opus-5.5 at 83.2, Grok-4.7 at 82.7

What it counts

The scores come from real people using agents in Arena's Agent Arena, mostly on jobs like coding and document analysis. Arena checks each session's trace for failures it can actually verify.

The one most people will care about is the unauthorized action, where an agent goes past what it was asked or allowed to do. Then there's deceptive completion, which is an agent telling you the job's done when it isn't. The third, false attribution, is narrower (the agent claims the user said or wanted something that the user's own evidence contradicts).

Arena says it took those definitions from the system cards OpenAI and Anthropic publish. Which is neat. The labs' own yardstick, held up against their models in the wild by somebody else.

When an agent deceives a user, or does something it wasn't permitted to do, it can have serious consequences.

From Arena Raises $200M Series B at $3.1B Valuation

How the labs came out

Arena's launch post on X gives the headline numbers. OpenAI has the best rate on every signal, from 0.89% for unauthorized actions to 2.34% for deceptive completion. Arena also says every lab's newer models beat their older ones.

Arena.ai@arena

Introducing the Arena Alignment Index, our new benchmark measuring safety and alignment of AI agents in real-world use. Built from 90K+ real-world agent sessions across 27 models, the index measures three critical signals: - Unauthorized Action (UA): Taking actions beyond the user's instructions or permissions - False Attribution (FA): Attributing statements or actions that are contradicted by user-provided evidence - Deceptive Completion (DC): Claiming a task was completed when it was not. Key findings: - OpenAI models currently lead the Alignment Index - Rogue actions are rare, but can have serious consequences when they occur - Agents can mislead users about task progress - Misalignment risks increase with conversation length - Safety and alignment are improving across model generations As shown in the leaderboard below (sorted by lab), @OpenAI’s GPT-6.1-Sol leads the Arena Alignment Index with a score of 87.9, followed by @AnthropicAI’s Claude-Opus-5.5 at 83.2 and @SpaceXAI's Grok-4.7 at 82.7. OpenAI also has the best observed rates across all three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. Across all four labs, newer models consistently outperform their predecessors, suggesting broad progress in agent safety and alignment. As agents take on longer, more complex, and higher-stakes tasks, measuring not just what they can accomplish, but how safely and reliably they act, becomes increasingly important. This marks an important step toward making safety and alignment a core part of how Arena evaluates AI. The index is an initial starting point, and we'll continue expanding the index with additional safety signals and models over time. More analysis below👇

View the post on X

I'd read past the ranking to two of the findings. Rogue actions are rare, Arena says, but they can do real damage. And the risk climbs as a conversation gets longer. So leave an agent working for an hour and it's more likely to wander somewhere you never sent it.

Why live sessions and not a fixed test? Arena's answer is that "static benchmarks break down once models recognize they're being tested."

What it doesn't cover yet

Arena calls this a preview. It "starts narrow", with three signals, and the company says it'll add more signals and models over time. The scores also only reflect what people happened to do in Agent Arena. A model that rarely gets long, risky jobs there (and Arena's own finding says length is where things go wrong) simply has fewer chances to fail.

The funding was co-led by Lightspeed Venture Partners and Khosla Ventures. Arena says it's passed $100 million in annualized revenue since raising its Series A in January 2026.

More on AI safety

All AI safety stories