Pull down to go back
Opus 4.7 dominates LLM Debate Benchmark with perfect record: 51 wins, 0 losses, crushing Sonnet 4.6 by 106 points

Opus 4.7 dominates LLM Debate Benchmark with perfect record: 51 wins, 0 losses, crushing Sonnet 4.6 by 106 points

Opus 4.7 在 LLM 辯論基準測試中完美無敗:51 勝 0 敗,領先 Sonnet 4.6 達 106 分

Anthropic's Opus 4.7 just achieved something remarkable on the LLM Debate Benchmark—it went undefeated. The model competed in 55 completed debates (arguing both sides of each motion), racking up 51 wins, 4 ties, and zero losses. That's a 106-point lead over the previous champion, Sonnet 4.6. Here's what makes this interesting: Opus 4.7 doesn't just win arguments—it wins by identifying the core issue (the "hinge") of each debate, then anchoring the entire discussion around it. This forces opponents to argue on Opus's terms, which is a genuinely clever rhetorical strategy. The benchmark uses a three-model judging panel to score each debate, and they deliberately avoid having judges from the same model family as the debaters (so no Opus judges for Opus debates). The full transcripts, model profiles, and detailed comparisons are available on GitHub if you want to see how these AI arguments actually play out.