Open-weight models (Llama, Qwen, DeepSeek, Mistral…) ranked side by side with the proprietary frontier (GPT, Claude, Gemini, Grok) — live pricing, context, benchmarks and capabilities. No marketing — just the facts.
| # | Model | Creator | Intelligence | Coding |
|---|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | Anthropic | 60.7 | 78.0 |
| 2 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 59.9 | 76.5 |
| 3 | GPT-5.6 Sol (max) | OpenAI | 58.9 | 77.4 |
| 4 | Kimi K3 | Kimi | 57.1 | 76.2 |
| 5 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | Anthropic | 55.7 | 74.3 |
| 6 | GPT-5.6 Terra (max) | OpenAI | 55.0 | 76.7 |
| 7 | GPT-5.5 (xhigh) | OpenAI | 54.8 | 74.9 |
| 8 | Grok 4.5 (high) | SpaceXAI | 53.8 | 72.4 |
| 9 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | Anthropic | 53.5 | 73.6 |
| 10 | Claude Sonnet 5 (Adaptive Reasoning, Max Effort) | Anthropic | 53.4 | 71.5 |
| 11 | GPT-5.4 (xhigh) | OpenAI | 51.4 | 71.1 |
| 12 | GPT-5.6 Luna (max) | OpenAI | 51.2 | 71.4 |
| 13 | GLM-5.2 (max) | Z AI | 51.1 | 68.8 |
| 14 | Muse Spark 1.1 (xhigh) | Meta | 50.6 | 71.3 |
| 15 | Gemini 3.5 Flash (high) | 50.2 | 70.1 | |
| 16 | Gemini 3.6 Flash (high) | 50.1 | 69.2 | |
| 17 | Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort) | Anthropic | 47.2 | 63.0 |
| 18 | Gemini 3.1 Pro Preview | 46.5 | 68.8 | |
| 19 | Qwen3.7 Max | Alibaba | 46.0 | 66.0 |
| 20 | MiniMax-M3 | MiniMax | 44.4 | 58.6 |
| 21 | DeepSeek V4 Pro (Reasoning, Max Effort) | DeepSeek | 44.3 | 59.4 |
| 22 | GPT-5.3 Codex (xhigh) | OpenAI | 44.3 | — |
| 23 | Kimi K2.6 | Kimi | 44.2 | 61.8 |
| 24 | Motif 3 (Beta) | Motif Technologies | 44.1 | 62.0 |
| 25 | Claude Opus 4.6 (Adaptive Reasoning, Max Effort) | Anthropic | 43.7 | — |
| 26 | Muse Spark | Meta | 43.1 | 58.6 |
| 27 | GPT-5.2 (xhigh) | OpenAI | 42.2 | — |
| 28 | MiMo-V2.5-Pro | Xiaomi | 42.2 | 60.2 |
| 29 | Kimi K2.7 Code | Kimi | 41.9 | 60.8 |
| 30 | Hy3 | Tencent | 41.2 | 58.8 |
| 31 | Nex-N2-Pro | Nex AGI | 41.0 | 59.1 |
| 32 | Claude Opus 4.5 (Reasoning) | Anthropic | 40.8 | — |
| 33 | Inkling (xhigh) | Thinking Machines | 40.7 | 52.1 |
| 34 | DeepSeek V4 Flash (Reasoning, Max Effort) | DeepSeek | 40.3 | 56.2 |
| 35 | MiMo-V2-Pro | Xiaomi | 40.3 | — |
| 36 | GLM-5.1 (Reasoning) | Z AI | 40.2 | 55.8 |
| 37 | GPT-5.2 Codex (xhigh) | OpenAI | 40.1 | — |
| 38 | GPT-5.4 mini (xhigh) | OpenAI | 40.0 | 56.1 |
| 39 | Qwen3.6 Max Preview | Alibaba | 40.0 | — |
| 40 | Grok Build 0.1 0616 | SpaceXAI | 39.8 | 51.5 |
| 41 | Gemini 3 Pro Preview (high) | 39.6 | — | |
| 42 | Qwen3.6 Plus | Alibaba | 39.6 | 54.5 |
| 43 | GLM-5 (Reasoning) | Z AI | 39.5 | — |
| 44 | Qwen3.7 Plus | Alibaba | 39.0 | 55.9 |
| 45 | Agnes 2.5 Pro Alpha | Sapiens AI | 38.8 | 58.8 |
| 46 | JT-4.1 Flash 236B A21B | China Mobile | 38.8 | 52.4 |
| 47 | GPT-5.4 nano (xhigh) | OpenAI | 38.2 | 56.1 |
| 48 | GLM-5-Turbo | Z AI | 38.1 | — |
| 49 | MiniMax-M2.7 | MiniMax | 38.1 | 52.6 |
| 50 | Gemini 3 Flash Preview (Reasoning) | 37.8 | — | |
| 51 | Nemotron 3 Ultra 550B A55B (Reasoning) | NVIDIA | 37.8 | 49.3 |
| 52 | Grok 4.3 (high) | SpaceXAI | 37.6 | 42.2 |
| 53 | MiMo-V2.5 | Xiaomi | 37.2 | 56.8 |
| 54 | Qwen3.6 27B (Reasoning) | Alibaba | 37.1 | 53.7 |
| 55 | Grok 4.20 0309 v2 (Reasoning) | SpaceXAI | 37.0 | — |
| 56 | GPT-5.1 (high) | OpenAI | 36.9 | 49.4 |
| 57 | Gemini 3.5 Flash-Lite | 36.5 | 49.3 | |
| 58 | Grok 4.20 0309 (Reasoning) | SpaceXAI | 36.5 | — |
| 59 | Claude 4.5 Sonnet (Reasoning) | Anthropic | 36.4 | 52.1 |
| 60 | MiMo-V2-Omni-0327 | Xiaomi | 36.4 | — |
How this is ranked: models are ordered by real usage popularity (weekly tokens processed across apps), via the OpenRouter API. Only open-weight models are included (Llama, DeepSeek, Qwen, Mistral, Gemma, Kimi, GLM, Phi, Nemotron, and more). Pricing is per million tokens; context is the maximum window. Movement arrows compare today's rank to the previous snapshot. Data refreshes daily.