AI model leaderboard 2026
46 leading large language models compared on context window, API price and best-fit use case. Updated July 10, 2026.
Quick answer
As of July 2, 2026, the leaderboard spans 46 models from 21 developers. Context windows reach 10M tokens; API prices range widely — see the table, or compare any two.
| # | Model | Bq score | Developer | Context | Input /1M | Output /1M | Best for | |
|---|---|---|---|---|---|---|---|---|
| 01 | Claude Fable 5 New | 98 | Anthropic | 1M tokens (128K max output) | $10.00 | $50.00 | The hardest multi-step agentic and coding work, frontier reasoning, and long-running autonomous tasks where capability matters most. | → |
| 02 | GPT-5.6 Sol | 96 | OpenAI | ~1.05M tokens (128K max output) | $5.00 | $30.00 | The highest-stakes reasoning, complex agents and frontier coding where capability outweighs cost. | → |
| 03 | GPT-5.5 | 95 | OpenAI | ~1.05M tokens (128K max output) | $5.00 | $30.00 | Highest-stakes reasoning, complex agents, and frontier coding tasks where capability outweighs cost. | → |
| 04 | Claude Opus 4.8 | 93 | Anthropic | 1M tokens | $5.00 | $25.00 | Complex coding agents, long-horizon tasks, and workloads needing top reliability and reasoning. | → |
| 05 | Gemini 3.1 Pro | 92 | 2M tokens | $2.00 (under 200K; $4.00 above) | $12.00 (under 200K; $18.00 above) | Long-document and multimodal workloads, RAG over huge corpora, and Google Cloud-native apps. | → | |
| 06 | Grok 4.5 | 91 | xAI | 500K tokens | $2.00 ($0.50 cached) | $6.00 | Opus-class reasoning and agentic coding with high token efficiency and real-time X and web data. | → |
| 07 | Kimi K3 New | 91 | Moonshot AI | 1M tokens | — | — | Open-weight frontier coding and long agentic sessions; full weights due July 27. | → |
| 08 | Claude Opus 4.7 New | 90 | Anthropic | 1M tokens | $5.00 | $25.00 | Production coding agents, long-horizon workflow automation, and document or finance agent work. | → |
| 09 | GPT-5.6 Terra | 89 | OpenAI | ~1.05M tokens (128K max output) | $2.50 | $15.00 | Most production apps wanting near-flagship quality at roughly half the flagship price. | → |
| 10 | Muse Spark 1.1 New | 89 | Meta | 1M tokens | $1.25 | $4.25 | Multimodal agentic reasoning at low cost via Meta's first paid API (US preview). | → |
| 11 | Grok 4.3 | 88 | xAI | 1M tokens | $1.25 ($0.20 cached) | $2.50 | Reasoning and agentic tasks needing current/real-time info and low output costs. | → |
| 12 | Inkling New | 88 | Thinking Machines | 1M tokens | — (open weights) | — (open weights) | Open-weight multimodal MoE (975B/41B active) you can self-host under Apache 2.0. | → |
| 13 | Claude Sonnet 5 New | 87 | Anthropic | 1M tokens | $2.00 | $10.00 | Running AI agents affordably, multi-step automation, coding assistants and high-scale chat where you want near-flagship quality at lower cost. | → |
| 14 | GPT-5.4 | 86 | OpenAI | ~1M tokens (922K input, 128K output) | $2.50 | $15.00 | Most production apps wanting near-flagship quality at roughly half the flagship price. | → |
| 15 | DeepSeek V4 Pro | 84 | DeepSeek | 1M tokens (384K max output) | $0.435 | $0.87 | Cost-conscious teams wanting frontier reasoning and coding at a fraction of US-lab prices. | → |
| 16 | MiMo-V2.5-Pro New | 84 | Xiaomi | 1M tokens | $0.435 | $0.87 | Agentic coding and long-running autonomous agents on a budget, including self-hosted deployments. | → |
| 17 | LongCat-2.0 | 84 | Meituan | 1M tokens (128K max output) | $0.30 | $1.20 | Cost-efficient near-frontier agentic coding and open-weight self-hosting. | → |
| 18 | GPT-Live-1 New | 84 | OpenAI | — | API coming soon | API coming soon | Natural full-duplex voice conversations; ChatGPT's default voice model. | → |
| 19 | MiMo-V2.5 New | 83 | Xiaomi | 1M tokens | $0.105 | $0.28 | High-volume agentic pipelines, long-document and multimodal processing where cost matters most. | → |
| 20 | Tencent Hy3 New | 83 | Tencent (Hunyuan) | 256K tokens | $0.063 | $0.21 | Cost-efficient agent pipelines, coding agents and high-volume production workloads. | → |
| 21 | Qwen3.7 Max Updated | 82 | Alibaba | 1M tokens (64K max output) | $1.25 | $3.75 | Long-context agentic workloads, multilingual apps, and Asia-Pacific deployments. | → |
| 22 | Kimi K2.7 Code New | 81 | Moonshot AI | 256K tokens | $0.95 | $4.00 | Repository-scale refactoring and long multi-turn agentic coding with an open-weight, cost-efficient model. | → |
| 23 | Z.ai GLM-5.2 New | 80 | Z.ai (Zhipu AI) | 1M tokens | $0.95 | $3.00 | Cost-efficient agentic coding and self-hosted deployments where open weights matter. | → |
| 24 | GPT-5.6 Luna | 80 | OpenAI | ~1.05M tokens (128K max output) | $1.00 | $6.00 | High-volume tasks like classification, extraction and chat where cost and speed matter most. | → |
| 25 | MiniMax M3 New | 79 | MiniMax | 1M tokens (min 512K) | $0.60 | $2.40 | Long-context coding agents, computer-use automation and high-volume tasks wanting frontier quality with open weights. | → |
| 26 | Doubao Seed 2.1 Pro New | 78 | ByteDance | 256K tokens | $0.83 | $4.17 | Complex coding and long-chain agent tasks at low cost, especially for teams already on Volcano Engine. | → |
| 27 | Gemini 3.5 Flash | 77 | 1M tokens | $1.50 | $9.00 | High-volume multimodal apps and agents needing speed with large context. | → | |
| 28 | Sakana Fugu Ultra New | 76 | Sakana AI | 1M tokens (128K max output; tier changes above 272K) | $5.00 | $30.00 | Hard multi-step reasoning, coding and code review, scientific/research tasks, and quality-critical agentic workflows. | → |
| 29 | Claude Sonnet 4.6 Updated | 75 | Anthropic | 1M tokens | $3.00 | $15.00 | Everyday coding, agents, and production apps wanting Claude quality without flagship pricing. | → |
| 30 | Laguna M.1 New | 73 | Poolside | 262K tokens (32K max output) | $0.20 | $0.40 | Agentic software engineering: codebase exploration, multi-file edits, test loops and CLI agents. | → |
| 31 | Kimi K2.6 | 72 | Moonshot AI | 262K tokens (64K max output) | $0.95 | $4.00 | Open-model coding agents and budget-conscious agentic pipelines. | → |
| 32 | Llama 4 Maverick | 70 | Meta | 1M tokens | ≈$0.15 | ≈$0.60 | Teams wanting an open, multimodal model they can self-host or run cheaply via multiple providers. | → |
| 33 | Nemotron 3 Ultra New | 70 | NVIDIA | 1M tokens | $0.50 | $2.20 | Long-running agent workflows, long-context processing and cost-efficient self-hosted deployments. | → |
| 34 | Step 3.7 Flash New | 70 | StepFun | 256K tokens | $0.20 | $1.15 | High-volume, cost-sensitive agents, search workflows and structured outputs where speed matters most. | → |
| 35 | gpt-oss-120b New | 69 | OpenAI | 131K tokens | ≈$0.03 | ≈$0.15 | Cost-sensitive high-volume workloads, on-prem or privacy-constrained deployments, and custom fine-tunes. | → |
| 36 | Mistral Large 3 | 68 | Mistral AI | 256K tokens | $0.50 | $1.50 | European/open-license deployments and value-focused production apps. | → |
| 37 | DeepSeek V4 Flash | 66 | DeepSeek | 1M tokens (384K max output) | $0.14 ($0.0028 cache hit) | $0.28 | High-volume, repetitive workloads and budget pipelines needing long context. | → |
| 38 | Doubao Seed 2.1 Turbo New | 65 | ByteDance | ≈256K tokens | $0.45 | $2.25 | High-volume, cost-sensitive coding and agent workloads needing near-flagship quality at low cost. | → |
| 39 | Grok 4.1 Fast | 64 | xAI | 2M tokens | $0.20 ($0.05 cached) | $0.50 | Large-context, high-volume agentic workloads where cost and context size dominate. | → |
| 40 | Gemini 3 Flash | 62 | 1M tokens | $0.50 | $3.00 | Cost-sensitive multimodal and long-context tasks at scale. | → | |
| 41 | GPT-5.4 mini | 60 | OpenAI | ~400K tokens | $0.75 | $4.50 | High-volume tasks like classification, extraction, and chat where cost and speed matter most. | → |
| 42 | Claude Haiku 4.5 | 58 | Anthropic | 200K tokens | $1.00 | $5.00 | Real-time chat, routing, classification, and high-throughput pipelines on a budget. | → |
| 43 | Amazon Nova 2 Pro New | 56 | Amazon | 256K tokens | $1.25 | $10.00 | AWS-centric enterprises wanting a frontier-class model inside Bedrock. | → |
| 44 | Llama 4 Scout | 54 | Meta | 10M tokens | ≈$0.08 | ≈$0.30 | Extreme long-context tasks (whole codebases, large document sets) on an open, cheap model. | → |
| 45 | Gemini 2.5 Flash-Lite New | 52 | 1M tokens (~64K output) | $0.10 | $0.40 | High-volume classification, summarization, extraction and simple chat at scale (note the Oct 2026 Gemini API retirement). | → | |
| 46 | Amazon Nova Premier Updated | 45 | Amazon | 1M tokens | $2.50 | $12.50 | AWS-native enterprises needing a managed multimodal model with strong governance. | → |
Ranked by the Benchquill score — a 0–100 capability-first editorial index weighing reasoning quality, agentic & coding strength, context window and price-performance (how we rank). Prices are public API list prices per 1M tokens (USD) and may change — verify with each provider. July 2, 2026.
Not sure which model to pick?
Compare any two models side-by-side on price, context and capability.
Compare models