What the perfect model actually looks like

Three good things pull against each other: smarter, faster, cheaper. Two go on the axes; the third becomes colour (green = better). Pick what matters and read the leaderboard — there is rarely a single winner.

X
Y
Colour

30 of 50 models have both selected axes. Hover for detail. Ringed dots are on the Pareto frontier; the is your current best pick.

Intelligence (higher is better)
GPT-5.5 (xhigh)
GPT-5.5 (high)
GPT-5.4
GPT-5.4 mini
gpt-oss-120b
Gemini 3.5 Flash
Gemini 3.1 Flash-Lite
Gemini 3.1 Pro
Gemma 4 31B
Claude Opus 4.8
Claude Opus 4.7
Claude Sonnet 4.6
Claude 4.5 Haiku
DeepSeek V4 Pro
DeepSeek V4 Flash
Qwen3.7 Max
Qwen3.7 Plus
Qwen3.6 Plus
Qwen3.6 35B
Qwen3.5 397B
GLM-5.2
GLM-5.1
Kimi K2.6
Kimi K2.7 Code
MiniMax-M3
MiniMax-M2.7
MiMo-V2.5-Pro
Grok 4.3
Nemotron 3 Ultra
Mistral Medium 3.5

Time / task (lower is better)

worsecolour = cost / task (grey = no data)better
Intelligence55%
Speed15%
Low cost30%

These weights rank the 30 models with complete intelligence + time + cost. The chart axes above are independent.

#ModelIntelTime$/tasktok/sScore
1GPT-5.5 (high)OpenAIfrontier53.13.36m$0.725966
2GPT-5.5 (xhigh)OpenAIfrontier54.84.81m$0.835866
3DeepSeek V4 ProDeepSeekfrontier44.38.23m$0.04847564
4DeepSeek V4 FlashDeepSeekfrontier40.36.78m$0.027910763
5Gemini 3.1 ProGooglefrontier46.51.74m$0.337212363
6Gemini 3.5 FlashGooglefrontier50.22.91m$0.681116163
7GLM-5.2Z AIfrontier51.16.01m$0.464510263
8MiMo-V2.5-ProXiaomifrontier42.28.88m$0.03224562
9GPT-5.4OpenAI51.43.87m$0.9915061
10Claude Opus 4.8Anthropicfrontier55.76.95m$2.056259
11MiniMax-M3MiniMaxfrontier44.46.51m$0.15676259
12Qwen3.7 PlusAlibabafrontier396.63m$0.04925058
13Claude Opus 4.7Anthropic53.56.2m$1.9757
14Qwen3.7 MaxAlibaba463.33m$0.689556
15Kimi K2.7 CodeKimifrontier41.95.25m$0.18385456
16Grok 4.3xAIfrontier37.61.44m$0.185815254
17MiniMax-M2.7MiniMax38.17.25m$0.07424453
18Nemotron 3 UltraNVIDIAfrontier37.82.36m$0.244617051
19GLM-5.1Z AI40.25.57m$0.24048251
20Qwen3.6 PlusAlibaba39.66.4m$0.266748
21Kimi K2.6Kimi42.811.68m$0.31464445
22Gemini 3.1 Flash-LiteGooglefrontier*251.14m$0.04328444
23Qwen3.6 35BAlibabafrontier31.62.64m$0.178417543
24GPT-5.4 miniOpenAI407.8m$0.504817343
25Claude Sonnet 4.6Anthropic47.213.2m$1.145842
26Gemma 4 31BGooglefrontier29.45.57m$0.083441
27Qwen3.5 397BAlibaba33.75.15m$0.33315139
28gpt-oss-120bOpenAI23.81.75m$0.060733839
29Claude 4.5 HaikuAnthropic312.79m$0.3311038
30Mistral Medium 3.5Mistral29.95.23m$0.59718130
Incomplete on Artificial Analysis (not ranked)
Claude Fable 5Anthropic60$3.25
GPT-5.5 (medium)OpenAI47
Gemini 3.5 Flash (medium)Google45159
Muse SparkMuse43
DeepSeek V4 Pro (high)DeepSeek4165
MiMo-V2-ProXiaomi40
GLM-5Z AI4075
GLM-5-TurboZ AI38
Kimi K2.5Kimi3853
DeepSeek V4 Flash (high)DeepSeek37
GLM-5V-TurboZ AI35
MiniMax-M2.5MiniMax34186
Hunyuan 3Tencent34119
DeepSeek V4 Pro (non-reasoning)DeepSeek3176
GPT-5 mini (medium)OpenAI31100
Grok 4.1 FastxAI31
DeepSeek V4 Flash (non-reasoning)DeepSeek29101
Trinity LargeTrinity25160
Gemini 3 FlashGoogle169
Grok 4.3 (medium)xAI155

All metrics per task, from Artificial Analysis (19 Jun 2026): Intelligence Index, time per task (min), cost per task (USD, AA’s official figure), output speed (tok/s) and output tokens per task. AA’s scrape caps each chart at 20 rows, so values for higher-cost/slower models (Opus, Fable, Sonnet, GPT-5.5 tiers) are read from AA’s own cost/time charts and marked *. A model is listed as incomplete only where AA shows no value at all. The Pareto frontier and ranking use the 30 fully-characterised models. Figures move — refresh modelData.ts before quoting.