What the perfect model actually looks like
Three good things pull against each other: smarter, faster, cheaper. Two go on the axes; the third becomes colour (green = better). Pick what matters and read the leaderboard — there is rarely a single winner.
30 of 50 models have both selected axes. Hover for detail. Ringed dots are on the Pareto frontier; the ★ is your current best pick.
Time / task (lower is better)
These weights rank the 30 models with complete intelligence + time + cost. The chart axes above are independent.
| # | Model | Intel | Time | $/task | tok/s | Score |
|---|---|---|---|---|---|---|
| 1 | GPT-5.5 (high)OpenAIfrontier | 53.1 | 3.36m | $0.72 | 59 | 66 |
| 2 | GPT-5.5 (xhigh)OpenAIfrontier | 54.8 | 4.81m | $0.83 | 58 | 66 |
| 3 | DeepSeek V4 ProDeepSeekfrontier | 44.3 | 8.23m | $0.0484 | 75 | 64 |
| 4 | DeepSeek V4 FlashDeepSeekfrontier | 40.3 | 6.78m | $0.0279 | 107 | 63 |
| 5 | Gemini 3.1 ProGooglefrontier | 46.5 | 1.74m | $0.3372 | 123 | 63 |
| 6 | Gemini 3.5 FlashGooglefrontier | 50.2 | 2.91m | $0.6811 | 161 | 63 |
| 7 | GLM-5.2Z AIfrontier | 51.1 | 6.01m | $0.4645 | 102 | 63 |
| 8 | MiMo-V2.5-ProXiaomifrontier | 42.2 | 8.88m | $0.0322 | 45 | 62 |
| 9 | GPT-5.4OpenAI | 51.4 | 3.87m | $0.99 | 150 | 61 |
| 10 | Claude Opus 4.8Anthropicfrontier | 55.7 | 6.95m | $2.05 | 62 | 59 |
| 11 | MiniMax-M3MiniMaxfrontier | 44.4 | 6.51m | $0.1567 | 62 | 59 |
| 12 | Qwen3.7 PlusAlibabafrontier | 39 | 6.63m | $0.0492 | 50 | 58 |
| 13 | Claude Opus 4.7Anthropic | 53.5 | 6.2m | $1.97 | — | 57 |
| 14 | Qwen3.7 MaxAlibaba | 46 | 3.33m | $0.68 | 95 | 56 |
| 15 | Kimi K2.7 CodeKimifrontier | 41.9 | 5.25m | $0.1838 | 54 | 56 |
| 16 | Grok 4.3xAIfrontier | 37.6 | 1.44m | $0.1858 | 152 | 54 |
| 17 | MiniMax-M2.7MiniMax | 38.1 | 7.25m | $0.0742 | 44 | 53 |
| 18 | Nemotron 3 UltraNVIDIAfrontier | 37.8 | 2.36m | $0.2446 | 170 | 51 |
| 19 | GLM-5.1Z AI | 40.2 | 5.57m | $0.2404 | 82 | 51 |
| 20 | Qwen3.6 PlusAlibaba | 39.6 | 6.4m | $0.2667 | — | 48 |
| 21 | Kimi K2.6Kimi | 42.8 | 11.68m | $0.3146 | 44 | 45 |
| 22 | Gemini 3.1 Flash-LiteGooglefrontier* | 25 | 1.14m | $0.043 | 284 | 44 |
| 23 | Qwen3.6 35BAlibabafrontier | 31.6 | 2.64m | $0.1784 | 175 | 43 |
| 24 | GPT-5.4 miniOpenAI | 40 | 7.8m | $0.5048 | 173 | 43 |
| 25 | Claude Sonnet 4.6Anthropic | 47.2 | 13.2m | $1.14 | 58 | 42 |
| 26 | Gemma 4 31BGooglefrontier | 29.4 | 5.57m | $0.08 | 34 | 41 |
| 27 | Qwen3.5 397BAlibaba | 33.7 | 5.15m | $0.3331 | 51 | 39 |
| 28 | gpt-oss-120bOpenAI | 23.8 | 1.75m | $0.0607 | 338 | 39 |
| 29 | Claude 4.5 HaikuAnthropic | 31 | 2.79m | $0.33 | 110 | 38 |
| 30 | Mistral Medium 3.5Mistral | 29.9 | 5.23m | $0.5971 | 81 | 30 |
| Incomplete on Artificial Analysis (not ranked) | ||||||
| — | Claude Fable 5Anthropic | 60 | — | $3.25 | — | — |
| — | GPT-5.5 (medium)OpenAI | 47 | — | — | — | — |
| — | Gemini 3.5 Flash (medium)Google | 45 | — | — | 159 | — |
| — | Muse SparkMuse | 43 | — | — | — | — |
| — | DeepSeek V4 Pro (high)DeepSeek | 41 | — | — | 65 | — |
| — | MiMo-V2-ProXiaomi | 40 | — | — | — | — |
| — | GLM-5Z AI | 40 | — | — | 75 | — |
| — | GLM-5-TurboZ AI | 38 | — | — | — | — |
| — | Kimi K2.5Kimi | 38 | — | — | 53 | — |
| — | DeepSeek V4 Flash (high)DeepSeek | 37 | — | — | — | — |
| — | GLM-5V-TurboZ AI | 35 | — | — | — | — |
| — | MiniMax-M2.5MiniMax | 34 | — | — | 186 | — |
| — | Hunyuan 3Tencent | 34 | — | — | 119 | — |
| — | DeepSeek V4 Pro (non-reasoning)DeepSeek | 31 | — | — | 76 | — |
| — | GPT-5 mini (medium)OpenAI | 31 | — | — | 100 | — |
| — | Grok 4.1 FastxAI | 31 | — | — | — | — |
| — | DeepSeek V4 Flash (non-reasoning)DeepSeek | 29 | — | — | 101 | — |
| — | Trinity LargeTrinity | 25 | — | — | 160 | — |
| — | Gemini 3 FlashGoogle | — | — | — | 169 | — |
| — | Grok 4.3 (medium)xAI | — | — | — | 155 | — |
All metrics per task, from Artificial Analysis (19 Jun 2026): Intelligence Index, time per task (min), cost per task (USD, AA’s official figure), output speed (tok/s) and output tokens per task. AA’s scrape caps each chart at 20 rows, so values for higher-cost/slower models (Opus, Fable, Sonnet, GPT-5.5 tiers) are read from AA’s own cost/time charts and marked *. A model is listed as incomplete only where AA shows no value at all. The Pareto frontier and ranking use the 30 fully-characterised models. Figures move — refresh modelData.ts before quoting.