Does an AI agent do what you asked, and only that? On October 8, 2026, Arena, the company behind a popular AI model leaderboard, launched the Arena Alignment Index to test exactly that. It scores 27 models using real sessions from its Agent Arena platform. OpenAI’s models hold the top five spots. But the index is new, its method has limits, and the company behind it just raised $200 million. Here is what it measures and what it cannot yet prove.
What the Index Measures
Arena looks for three failures in real agent conversations. Unauthorized action means the agent did something beyond what the user asked or was allowed. False attribution means the agent credits the user with a statement, request, or fact that the user’s own evidence contradicts. Deceptive completion means the agent says a task is done when evidence shows it is not.
Arena wrote rubrics for each failure and had an AI judge apply them to sampled sessions. A session counts only if the judge can point to a specific claim or action and the evidence for it. Humans reviewed the rubrics, and Arena changed them where reviewers disagreed with the judge. It then adjusted the rates for conversation length.
Each failure rate becomes a score of one minus its square root, which makes small gaps near the top easier to see. The three scores are averaged, with unauthorized action counting for half and the other two for a quarter each. Those weights are Arena’s own choice. Arena also says the three signals cover only a small part of safety and alignment.
What the Scores Show
GPT-6.1 Sol scores 87.9, Claude Opus 5.5 scores 83.2, and Grok 4.7 scores 82.7. Opus 5.5 ranks sixth and Grok 4.7 seventh, because OpenAI’s models fill the top five. Four of them score between 87.6 and 87.9.
Look at the error margins. The top score is 87.9, plus or minus 1.5. The next two are 87.8, plus or minus 1.1 and 1.4. Those ranges overlap, so first place is not clearly better than second, third, or fourth. Arena’s own rank ranges show the same.
The failures themselves are rare but real. Arena reports an unauthorized action in about 0.9% of GPT-6.1 Sol’s sessions and 1.25% of Opus 5.5’s. Deceptive completion is more common, at 2.3% and 6.4%. Across all models, about 10% of sessions had a deceptive completion, and in code debugging that rose to 48%. Longer chats fail more often. In sessions with 20 or more messages, about one in eight had an unauthorized action.
The Chinese open models Arena tested score lower, between 69 and 75. Kimi K3 scores 75.0 and GLM 5.3 scores 75.1.
What the Index Cannot Tell You
This is Arena’s own method and data. No outside group has checked the rubrics or the AI judge’s calls, and the blog I read does not name the judge model. Arena has published research saying AI judges prefer their own answers about 70% more often than humans do. That does not show bias here, but it is a fair question.
The data comes from people who choose to use Agent Arena, which may not match your work. Arena’s leaderboard page shows 72,509 sessions, while its blog says 90,000. I could not explain the gap.
Deceptive completion also depends partly on skill. Arena says weaker models may fail to deliver what they promise, so a high rate can mean less capable, not more dishonest. And the breakdown of how models fail rests on small numbers. For Opus 5.5, the percentages come in tenths, which suggests about 10 flagged cases. That is my inference.
The Business Context
Arena began as a Berkeley research project. On the same day as the index, it said it raised $200 million at a $3.1 billion valuation, led by Lightspeed and Khosla Ventures. TechCrunch reports that Arena also sells evaluation services to AI labs and companies. Arena calls itself a neutral third party. I found nothing showing that this affected the scores. Still, a score like this can easily end up in launch marketing, so outside checks matter.
What Could Come Next
In the near term, Arena says it will add more signals, starting with how well models refuse harmful prompts, and more models. Outside researchers could test the rubrics on their own data.
The long-term picture is speculation. If indexes like this become standard, buyers may choose agents by their failure patterns, not only by price and performance, as Arena suggests. Whether this one earns that trust depends on whether others can check it.