Best agent model

Long-running tasks with tool calls, where being nearly right at step three ruins step nine.

Right now

Claude Opus 5 (High) from Anthropic currently ranks first for multi-step work with tools, ahead of Claude Opus 5 (Max) from Anthropic. The ranking comes from LMArena and was last published 2026-08-31.

Leader
Claude Opus 5 (High)
Made by
Anthropic
Score
13.8
Source
LMArena

Full ranking Human head-to-head preference votes, Bradley-Terry rating.

Rank# Model Made by Score % Votes
1 Claude Opus 5 (High) Anthropic 13.8 21,433
2 Claude Opus 5 (Max) Anthropic 11.6 17,073
3 Claude Fable 5 (High) Anthropic 10.6 34,905
4 GPT 5.6 Sol (xHigh) OpenAI 9.8 28,524
5 Claude Opus 4.8 (High) Anthropic 9.5 36,887
6 Kimi K3 (Max) Open Moonshot AI 8.7 94,549
7 GPT 5.5 (xHigh) OpenAI 7.9 50,040
8 Claude Sonnet 5 (High) Anthropic 7.5 27,467
9 Claude Opus 4.7 (High) Anthropic 6.6 36,570
10 Claude Opus 4.7 Anthropic 6.3 37,121
11 GPT 5.5 (High) OpenAI 6.1 73,880
12 GLM 5.2 (Max) Open Zai 6.1 65,079
13 Grok 4.5 xAI 6.1 34,008
14 Qwen3.8 Max Alibaba 6.0 18,386
15 DeepSeek V4 Pro (High) (0813) Open DeepSeek 5.9 23,928
16 Grok 4.6 (xHigh) xAI 5.8 15,090
17 Claude Opus 4.6 Anthropic 5.2 36,365
18 GPT 5.5 OpenAI 4.8 77,762
19 GLM 5.3 Flash Open Zai 4.4 9,970
20 GLM 5.3 (Max) Open Zai 3.8 41,022
21 GPT 5.4 (High) OpenAI 3.2 76,973
22 Deepseek V4 Flash (High) (20260731) Open DeepSeek 3.0 51,219
23 GPT 5.6 Terra (xHigh) OpenAI 2.9 16,778
24 Qwen3.8 Flash Next Open Alibaba 2.4 8,777
25 Claude Opus 4.8 Anthropic 2.3 33,690

Published by LMArena on 2026-08-31 · copied here 2026-09-02 · no adjustment applied

Other leaderboards