Leaderboards
Which model is best, for what.
One ranking per job, not one ranking for everything. Rebuilt every day from data the benchmark operators publish themselves — each board names its source and the date that source last updated.
Best overall chat model
Head-to-head human preference across everyday prompts. The closest thing there is to a general ranking.
| Rank# | Model | Made by | Arena score | Votes |
|---|---|---|---|---|
| 1 | Claude Opus 5 Max | Anthropic | 1,505 | 16,839 |
| 2 | Claude Opus 5 High | Anthropic | 1,504 | 34,617 |
| 3 | Claude Opus 4.6 High | Anthropic | 1,503 | 72,110 |
| 4 | Claude Opus 4.6 | Anthropic | 1,498 | 76,032 |
| 5 | Claude Fable 5 | Anthropic | 1,494 | 26,977 |
| 6 | Gemini 3.7 Flash High | 1,491 | 5,685 | |
| 7 | Claude Opus 4.7 High | Anthropic | 1,490 | 60,123 |
| 8 | muse-spark-1.2 (xHigh) | Meta | 1,488 | 3,244 |
| 9 | Gemini 3.5 Flash High | 1,484 | 33,874 | |
| 10 | Claude Opus 4.7 | Anthropic | 1,483 | 61,256 |
| 11 | Gemini 3.1 Pro Preview | 1,480 | 102,763 | |
| 12 | Qwen3.8 Max | Alibaba | 1,480 | 13,136 |
| 13 | Gemini 3 Pro | 1,480 | 40,675 | |
| 14 | Muse Spark 1.1 | Meta | 1,479 | 23,768 |
| 15 | Gemini 3.6 Flash High | 1,476 | 22,009 | |
| 16 | Kimi K3 Max Open | Moonshot AI | 1,476 | 17,895 |
| 17 | Gemini 3.5 Flash Medium | 1,476 | 32,274 | |
| 18 | Glm 5.3 Max Open | Zai | 1,474 | 7,401 |
| 19 | Qwen3.7 Max Preview | Alibaba | 1,474 | 3,710 |
| 20 | Muse Spark | Meta | 1,474 | 13,571 |
| 21 | GPT 5.5 High | OpenAI | 1,472 | 63,265 |
| 22 | Qwen3.5 Max Preview | Alibaba | 1,471 | 21,487 |
| 23 | Glm 5.3 Flash Open | Zai | 1,471 | 4,399 |
| 24 | GPT 5.4 High | OpenAI | 1,470 | 60,657 |
| 25 | Ernie 5.1 | Baidu | 1,468 | 37,137 |
Best model for coding
Judged on real build tasks where a human picks the better result, not on a self-reported pass rate.
| Rank# | Model | Made by | Arena score | Votes |
|---|---|---|---|---|
| 1 | qwen3.8-max-0902 | Alibaba | 1,691 | 1,389 |
| 2 | Claude Opus 5 Max | Anthropic | 1,688 | 10,334 |
| 3 | Kimi K3 Max Open | Moonshot AI | 1,674 | 4,544 |
| 4 | qwen3.8-max | Alibaba | 1,669 | 3,219 |
| 5 | Claude Opus 5 High | Anthropic | 1,661 | 10,326 |
| 6 | Hy4 Preview Open | Tencent | 1,629 | 1,368 |
| 7 | Grok 4.6 High | xAI | 1,629 | 1,532 |
| 8 | Claude Fable 5 | Anthropic | 1,628 | 9,066 |
| 9 | Qwen3.8 Flash Next Open | Alibaba | 1,620 | 1,876 |
| 10 | gpt-5.6-sol-xhigh (codex-harness) | OpenAI | 1,616 | 10,684 |
| 11 | Glm 5.3 Max Open | Zai | 1,608 | 2,648 |
| 12 | Glm 5.3 Flash Open | Zai | 1,604 | 1,698 |
| 13 | Qwen3.8 27b Open | Alibaba | 1,599 | 3,256 |
| 14 | Gemini 3.7 Flash High | 1,587 | 3,007 | |
| 15 | Glm 5.2 Max Open | Zai | 1,585 | 9,569 |
| 16 | Deepseek v4 Pro High Open | DeepSeek | 1,584 | 3,289 |
| 17 | Deepseek v4 Flash High Open | DeepSeek | 1,581 | 3,928 |
| 18 | Claude Opus 4.8 High | Anthropic | 1,563 | 12,603 |
| 19 | Claude Opus 4.7 | Anthropic | 1,557 | 15,430 |
| 20 | Claude Opus 4.7 High | Anthropic | 1,556 | 15,909 |
| 21 | Grok 4.5 | xAI | 1,555 | 7,126 |
| 22 | Claude Opus 4.6 High | Anthropic | 1,546 | 17,908 |
| 23 | Claude Opus 4.8 | Anthropic | 1,540 | 11,543 |
| 24 | Muse Spark 1.1 | Meta | 1,540 | 7,091 |
| 25 | Gemini 3.6 Flash High | 1,538 | 7,726 |
Best agent model
Long-running tasks with tool calls, where being nearly right at step three ruins step nine.
| Rank# | Model | Made by | Score % | Votes |
|---|---|---|---|---|
| 1 | Claude Opus 5 (High) | Anthropic | 13.8 | 21,433 |
| 2 | Claude Opus 5 (Max) | Anthropic | 11.6 | 17,073 |
| 3 | Claude Fable 5 (High) | Anthropic | 10.6 | 34,905 |
| 4 | GPT 5.6 Sol (xHigh) | OpenAI | 9.8 | 28,524 |
| 5 | Claude Opus 4.8 (High) | Anthropic | 9.5 | 36,887 |
| 6 | Kimi K3 (Max) Open | Moonshot AI | 8.7 | 94,549 |
| 7 | GPT 5.5 (xHigh) | OpenAI | 7.9 | 50,040 |
| 8 | Claude Sonnet 5 (High) | Anthropic | 7.5 | 27,467 |
| 9 | Claude Opus 4.7 (High) | Anthropic | 6.6 | 36,570 |
| 10 | Claude Opus 4.7 | Anthropic | 6.3 | 37,121 |
| 11 | GPT 5.5 (High) | OpenAI | 6.1 | 73,880 |
| 12 | GLM 5.2 (Max) Open | Zai | 6.1 | 65,079 |
| 13 | Grok 4.5 | xAI | 6.1 | 34,008 |
| 14 | Qwen3.8 Max | Alibaba | 6.0 | 18,386 |
| 15 | DeepSeek V4 Pro (High) (0813) Open | DeepSeek | 5.9 | 23,928 |
| 16 | Grok 4.6 (xHigh) | xAI | 5.8 | 15,090 |
| 17 | Claude Opus 4.6 | Anthropic | 5.2 | 36,365 |
| 18 | GPT 5.5 | OpenAI | 4.8 | 77,762 |
| 19 | GLM 5.3 Flash Open | Zai | 4.4 | 9,970 |
| 20 | GLM 5.3 (Max) Open | Zai | 3.8 | 41,022 |
| 21 | GPT 5.4 (High) | OpenAI | 3.2 | 76,973 |
| 22 | Deepseek V4 Flash (High) (20260731) Open | DeepSeek | 3.0 | 51,219 |
| 23 | GPT 5.6 Terra (xHigh) | OpenAI | 2.9 | 16,778 |
| 24 | Qwen3.8 Flash Next Open | Alibaba | 2.4 | 8,777 |
| 25 | Claude Opus 4.8 | Anthropic | 2.3 | 33,690 |
Best model for images and documents
Understanding what is in a picture, a chart or a scanned page.
| Rank# | Model | Made by | Arena score | Votes |
|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 1,329 | 10,002 |
| 2 | Claude Opus 5 High | Anthropic | 1,322 | 8,271 |
| 3 | Claude Opus 4.7 | Anthropic | 1,317 | 21,457 |
| 4 | Claude Opus 4.7 High | Anthropic | 1,316 | 21,137 |
| 5 | Claude Opus 4.6 High | Anthropic | 1,315 | 20,883 |
| 6 | Qwen3.8 Max | Alibaba | 1,313 | 7,244 |
| 7 | Claude Opus 4.6 | Anthropic | 1,311 | 25,210 |
| 8 | Gemini 3.5 Flash High | 1,311 | 8,259 | |
| 9 | Muse Spark | Meta | 1,306 | 5,585 |
| 10 | Gemini 3.5 Flash Medium | 1,306 | 8,261 | |
| 11 | Gemini 3 Pro | 1,304 | 13,019 | |
| 12 | muse-spark-1.2 (xHigh) | Meta | 1,304 | 1,844 |
| 13 | Gemini 3.6 Flash High | 1,302 | 3,816 | |
| 14 | GPT 5.4 High | OpenAI | 1,298 | 23,263 |
| 15 | Glm 5.3 Flash Open | Zai | 1,296 | 1,389 |
| 16 | GPT 5.5 | OpenAI | 1,296 | 21,389 |
| 17 | Gemini 3.1 Pro Preview | 1,295 | 39,412 | |
| 18 | GPT 5.5 High | OpenAI | 1,294 | 20,121 |
| 19 | Claude Opus 4.8 High | Anthropic | 1,294 | 13,521 |
| 20 | GPT 5.4 | OpenAI | 1,293 | 21,245 |
| 21 | Muse Spark 1.1 | Meta | 1,293 | 6,769 |
| 22 | Grok 4.5 | xAI | 1,291 | 6,401 |
| 23 | Claude Opus 4.8 | Anthropic | 1,289 | 13,980 |
| 24 | Gemini 3 Flash | 1,285 | 36,747 | |
| 25 | Claude Sonnet 4.6 | Anthropic | 1,282 | 25,610 |
Best model for search
Answers that cite the live web, ranked on whether the citation holds up.
| Rank# | Model | Made by | Arena score | Votes |
|---|---|---|---|---|
| 1 | GPT 5.6 Sol Xhigh | OpenAI | 1,257 | 29,663 |
| 2 | Claude Opus 4.6 Search | Anthropic | 1,253 | 134,699 |
| 3 | GPT 5.5 Search | OpenAI | 1,242 | 89,873 |
| 4 | Claude Opus 4.7 | Anthropic | 1,233 | 91,394 |
| 5 | Claude Fable 5 | Anthropic | 1,230 | 41,795 |
| 6 | Ernie 5.1 | Baidu | 1,227 | 3,788 |
| 7 | Claude Sonnet 4.6 Search | Anthropic | 1,221 | 134,905 |
| 8 | Grok 4.5 | xAI | 1,213 | 31,505 |
| 9 | Gemini 3.1 Pro Grounding | 1,210 | 113,282 | |
| 10 | Gemini 3 Pro Grounding | 1,207 | 37,024 | |
| 11 | GPT 5.2 Search | OpenAI | 1,207 | 52,712 |
| 12 | Claude Opus 4.8 | Anthropic | 1,204 | 70,998 |
| 13 | Grok 4.20 Multi Agent Beta | xAI | 1,204 | 109,553 |
| 14 | GPT 5.1 Search | OpenAI | 1,199 | 59,909 |
| 15 | Gemini 3 Flash Grounding | 1,198 | 149,334 | |
| 16 | GPT 5.4 Search | OpenAI | 1,197 | 110,116 |
| 17 | Claude Sonnet 5 Search | Anthropic | 1,194 | 40,230 |
| 18 | Grok 4.20 Beta1 | xAI | 1,189 | 53,921 |
| 19 | Claude Opus 4.5 Search | Anthropic | 1,180 | 61,573 |
| 20 | GPT 5.2 Search Non Reasoning | OpenAI | 1,172 | 75,658 |
| 21 | Grok 4.1 Fast Search | xAI | 1,171 | 81,507 |
| 22 | Grok 4 Fast Search | xAI | 1,171 | 41,794 |
| 23 | Grok 4.3 | xAI | 1,165 | 91,483 |
| 24 | Claude Sonnet 4.5 Search | Anthropic | 1,158 | 127,378 |
| 25 | Claude Opus 4.1 Search | Anthropic | 1,148 | 76,933 |
Best image generator
Text-to-image, ranked by people choosing between two results.
| Rank# | Model | Made by | Arena score | Votes |
|---|---|---|---|---|
| 1 | gpt-image-2 (medium) | OpenAI | 1,380 | 69,194 |
| 2 | Reve 2.1 | Reve | 1,302 | 7,598 |
| 3 | Muse Image | Meta | 1,283 | 14,510 |
| 4 | Reve 2.0 | Reve | 1,270 | 14,650 |
| 5 | gemini-3.1-flash-image (nano-banana-2) [web-search] | 1,263 | 27,910 | |
| 6 | Qwen Image 3.0 Pro | Alibaba | 1,258 | 3,735 |
| 7 | Seedream 5.0 Pro | ByteDance | 1,257 | 26,510 |
| 8 | Mai Image 2.5 | Microsoft-Ai | 1,256 | 44,940 |
| 9 | gemini-3.1-flash-lite-image (nano-banana-2-lite) | 1,251 | 16,419 | |
| 10 | gemini-3-pro-image-2k (nano-banana-pro) | 1,246 | 138,756 | |
| 11 | GPT Image 1.5 High Fidelity | OpenAI | 1,239 | 142,427 |
| 12 | gemini-3-pro-image-preview (nano-banana-pro) | 1,232 | 82,835 | |
| 13 | Grok Imagine Image Quality | xAI | 1,228 | 48,786 |
| 14 | Ideogram 4.0 Quality Open | Ideogram | 1,206 | 27,891 |
| 15 | Qwen Image 2.0 Pro 2026.06.22 | Alibaba | 1,191 | 12,026 |
| 16 | Uni 1.1 Max | Luma AI | 1,188 | 13,491 |
| 17 | Mai Image 2 | Microsoft-Ai | 1,182 | 49,207 |
| 18 | Cosmos3-Super-Text2Image (Agentic) Open | NVIDIA | 1,181 | 2,957 |
| 19 | Uni 1.1 | Luma AI | 1,180 | 26,927 |
| 20 | Grok Imagine Image | xAI | 1,172 | 217,461 |
| 21 | Recraft V4.1 Utility Pro | Recraft | 1,169 | 2,519 |
| 22 | Flux 2 Max | Bfl | 1,162 | 117,464 |
| 23 | Grok Imagine Image Pro | xAI | 1,161 | 93,711 |
| 24 | Cosmos3-Super-Text2Image Open | NVIDIA | 1,158 | 4,669 |
| 25 | Flux 2 Flex | Bfl | 1,156 | 149,402 |
Best video generator
Text-to-video, same method: two clips, a human picks one.
| Rank# | Model | Made by | Arena score | Votes |
|---|---|---|---|---|
| 1 | Gemini Omni 1.1 Flash | 1,515 | 1,762 | |
| 2 | Gemini Omni Flash | 1,512 | 19,830 | |
| 3 | Flux 3 Video | Bfl | 1,495 | 1,287 |
| 4 | Dreamina Seedance 2.0 720p | ByteDance | 1,479 | 51,266 |
| 5 | Dreamina Seedance 2.5 720p | ByteDance | 1,476 | 2,121 |
| 6 | Minimax H3 Open | MiniMax | 1,460 | 5,813 |
| 7 | Muse Video | Meta | 1,457 | 2,178 |
| 8 | Happyhorse 1.0 | Aorizon | 1,428 | 22,112 |
| 9 | Sora 2 Pro | OpenAI | 1,365 | 48,782 |
| 10 | Veo 3.1 Audio | 1,364 | 13,705 | |
| 11 | Veo 3.1 Audio 1080p | 1,363 | 24,972 | |
| 12 | Veo 3.1 Fast Audio | 1,361 | 39,371 | |
| 13 | Veo 3.1 Fast Audio 1080p | 1,358 | 25,768 | |
| 14 | Veo 3 Fast Audio | 1,347 | 25,169 | |
| 15 | Wan2.7 T2v | Wan | 1,344 | 20,882 |
| 16 | Grok Imagine Video 720p | xAI | 1,344 | 157,732 |
| 17 | Sora 2 | OpenAI | 1,341 | 65,433 |
| 18 | Veo 3 Audio | 1,340 | 18,941 | |
| 19 | Wan2.6 T2v | Alibaba | 1,329 | 47,267 |
| 20 | Seedance V1.5 Pro | ByteDance | 1,256 | 75,890 |
| 21 | Veo 3 | 1,253 | 14,951 | |
| 22 | Veo 3 Fast | 1,248 | 15,225 | |
| 23 | Wan2.5 T2v Preview | Alibaba | 1,247 | 30,497 |
| 24 | Pixverse V5.6 | Unknown | 1,240 | 32,296 |
| 25 | Runway Gen 4.5 | Runway | 1,224 | 41,583 |
Cheapest capable models
Live list prices across providers. Ranking here is arithmetic, not opinion.
| Rank# | Model | Made by | Per 1M tokens | Context |
|---|---|---|---|---|
| 1 | Mistral Nemo | Mistral | $0.022 | 131K |
| 2 | Ling-3.0-flash | — | $0.032 | 262K |
| 3 | Granite 4.0 Micro | IBM | $0.041 | 131K |
| 4 | Nex-N2-Mini | Nex AGI | $0.044 | 262K |
| 5 | Solar Pro 4 | Upstage | $0.052 | 524K |
| 6 | Qwen3.7 Flash | Qwen | $0.055 | 1M |
| 7 | gpt-oss-20b | OpenAI | $0.055 | 131K |
| 8 | Llama 3.1 8B Instruct | Meta | $0.057 | 131K |
| 9 | Nova Micro 1.0 | Amazon | $0.061 | 128K |
| 10 | Granite 4.1 8B | IBM | $0.062 | 131K |
| 11 | Gemma 3 4B | $0.062 | 131K | |
| 12 | Command R7B (12-2024) | Cohere | $0.066 | 128K |
| 13 | Mercury 2.5 Preview | Inception | $0.068 | 260K |
| 14 | GPT-5 Nano (batch) | OpenAI | $0.069 | 400K |
| 15 | gpt-oss-120b | OpenAI | $0.070 | 131K |
| 16 | Laguna XS 2.1 | Poolside | $0.075 | 262K |
| 17 | Gemma 3 12B | $0.075 | 131K | |
| 18 | DeepSeek V4 Flash Latest | — | $0.077 | 1.311M |
| 19 | Qwen3 30B A3B Instruct 2507 | Qwen | $0.084 | 262K |
| 20 | Nemotron 3 Nano 30B A3B | NVIDIA | $0.087 | 262K |
| 21 | gpt-oss-20b (batch) | OpenAI | $0.087 | 131K |
| 22 | Gemini 2.5 Flash Lite (batch) | $0.087 | 1.049M | |
| 23 | GPT-4.1 Nano (batch) | OpenAI | $0.087 | 1.048M |
| 24 | DeepSeek V4 Flash 0731 | DeepSeek | $0.094 | 1.311M |
| 25 | DeepSeek V4 Flash 0423 | DeepSeek | $0.099 | 1.049M |
How to read these
Arena scores are relative, not absolute. They come from people comparing two anonymous answers and picking one. A 20-point gap near the top is often inside the margin of error — treat the top few as a tie unless the confidence intervals separate.
A benchmark is not your workload. The board tells you which model wins on average across other people's prompts. Yours are different. Use it to pick two candidates, then test both on your own work.
Prices move faster than rankings. The price board is list price per million tokens, blended three parts input to one part output. Providers discount, cache and rate-limit differently, so it is a starting point, not a bill.
We do not run these tests. Every number here is published by the operator named next to it and copied without adjustment. If a board looks stale, its source has not updated — the date shown is theirs, not ours.