Leaderboard

AI model leaderboard 2026

46 leading large language models compared on context window, API price and best-fit use case. Updated July 10, 2026.

Quick answer

As of July 2, 2026, the leaderboard spans 46 models from 21 developers. Context windows reach 10M tokens; API prices range widely — see the table, or compare any two.

#ModelBq scoreDeveloperContextInput /1MOutput /1MBest for
01 Claude Fable 5 New 98 Anthropic 1M tokens (128K max output) $10.00 $50.00 The hardest multi-step agentic and coding work, frontier reasoning, and long-running autonomous tasks where capability matters most.
02 GPT-5.6 Sol 96 OpenAI ~1.05M tokens (128K max output) $5.00 $30.00 The highest-stakes reasoning, complex agents and frontier coding where capability outweighs cost.
03 GPT-5.5 95 OpenAI ~1.05M tokens (128K max output) $5.00 $30.00 Highest-stakes reasoning, complex agents, and frontier coding tasks where capability outweighs cost.
04 Claude Opus 4.8 93 Anthropic 1M tokens $5.00 $25.00 Complex coding agents, long-horizon tasks, and workloads needing top reliability and reasoning.
05 Gemini 3.1 Pro 92 Google 2M tokens $2.00 (under 200K; $4.00 above) $12.00 (under 200K; $18.00 above) Long-document and multimodal workloads, RAG over huge corpora, and Google Cloud-native apps.
06 Grok 4.5 91 xAI 500K tokens $2.00 ($0.50 cached) $6.00 Opus-class reasoning and agentic coding with high token efficiency and real-time X and web data.
07 Kimi K3 New 91 Moonshot AI 1M tokens Open-weight frontier coding and long agentic sessions; full weights due July 27.
08 Claude Opus 4.7 New 90 Anthropic 1M tokens $5.00 $25.00 Production coding agents, long-horizon workflow automation, and document or finance agent work.
09 GPT-5.6 Terra 89 OpenAI ~1.05M tokens (128K max output) $2.50 $15.00 Most production apps wanting near-flagship quality at roughly half the flagship price.
10 Muse Spark 1.1 New 89 Meta 1M tokens $1.25 $4.25 Multimodal agentic reasoning at low cost via Meta's first paid API (US preview).
11 Grok 4.3 88 xAI 1M tokens $1.25 ($0.20 cached) $2.50 Reasoning and agentic tasks needing current/real-time info and low output costs.
12 Inkling New 88 Thinking Machines 1M tokens — (open weights) — (open weights) Open-weight multimodal MoE (975B/41B active) you can self-host under Apache 2.0.
13 Claude Sonnet 5 New 87 Anthropic 1M tokens $2.00 $10.00 Running AI agents affordably, multi-step automation, coding assistants and high-scale chat where you want near-flagship quality at lower cost.
14 GPT-5.4 86 OpenAI ~1M tokens (922K input, 128K output) $2.50 $15.00 Most production apps wanting near-flagship quality at roughly half the flagship price.
15 DeepSeek V4 Pro 84 DeepSeek 1M tokens (384K max output) $0.435 $0.87 Cost-conscious teams wanting frontier reasoning and coding at a fraction of US-lab prices.
16 MiMo-V2.5-Pro New 84 Xiaomi 1M tokens $0.435 $0.87 Agentic coding and long-running autonomous agents on a budget, including self-hosted deployments.
17 LongCat-2.0 84 Meituan 1M tokens (128K max output) $0.30 $1.20 Cost-efficient near-frontier agentic coding and open-weight self-hosting.
18 GPT-Live-1 New 84 OpenAI API coming soon API coming soon Natural full-duplex voice conversations; ChatGPT's default voice model.
19 MiMo-V2.5 New 83 Xiaomi 1M tokens $0.105 $0.28 High-volume agentic pipelines, long-document and multimodal processing where cost matters most.
20 Tencent Hy3 New 83 Tencent (Hunyuan) 256K tokens $0.063 $0.21 Cost-efficient agent pipelines, coding agents and high-volume production workloads.
21 Qwen3.7 Max Updated 82 Alibaba 1M tokens (64K max output) $1.25 $3.75 Long-context agentic workloads, multilingual apps, and Asia-Pacific deployments.
22 Kimi K2.7 Code New 81 Moonshot AI 256K tokens $0.95 $4.00 Repository-scale refactoring and long multi-turn agentic coding with an open-weight, cost-efficient model.
23 Z.ai GLM-5.2 New 80 Z.ai (Zhipu AI) 1M tokens $0.95 $3.00 Cost-efficient agentic coding and self-hosted deployments where open weights matter.
24 GPT-5.6 Luna 80 OpenAI ~1.05M tokens (128K max output) $1.00 $6.00 High-volume tasks like classification, extraction and chat where cost and speed matter most.
25 MiniMax M3 New 79 MiniMax 1M tokens (min 512K) $0.60 $2.40 Long-context coding agents, computer-use automation and high-volume tasks wanting frontier quality with open weights.
26 Doubao Seed 2.1 Pro New 78 ByteDance 256K tokens $0.83 $4.17 Complex coding and long-chain agent tasks at low cost, especially for teams already on Volcano Engine.
27 Gemini 3.5 Flash 77 Google 1M tokens $1.50 $9.00 High-volume multimodal apps and agents needing speed with large context.
28 Sakana Fugu Ultra New 76 Sakana AI 1M tokens (128K max output; tier changes above 272K) $5.00 $30.00 Hard multi-step reasoning, coding and code review, scientific/research tasks, and quality-critical agentic workflows.
29 Claude Sonnet 4.6 Updated 75 Anthropic 1M tokens $3.00 $15.00 Everyday coding, agents, and production apps wanting Claude quality without flagship pricing.
30 Laguna M.1 New 73 Poolside 262K tokens (32K max output) $0.20 $0.40 Agentic software engineering: codebase exploration, multi-file edits, test loops and CLI agents.
31 Kimi K2.6 72 Moonshot AI 262K tokens (64K max output) $0.95 $4.00 Open-model coding agents and budget-conscious agentic pipelines.
32 Llama 4 Maverick 70 Meta 1M tokens ≈$0.15 ≈$0.60 Teams wanting an open, multimodal model they can self-host or run cheaply via multiple providers.
33 Nemotron 3 Ultra New 70 NVIDIA 1M tokens $0.50 $2.20 Long-running agent workflows, long-context processing and cost-efficient self-hosted deployments.
34 Step 3.7 Flash New 70 StepFun 256K tokens $0.20 $1.15 High-volume, cost-sensitive agents, search workflows and structured outputs where speed matters most.
35 gpt-oss-120b New 69 OpenAI 131K tokens ≈$0.03 ≈$0.15 Cost-sensitive high-volume workloads, on-prem or privacy-constrained deployments, and custom fine-tunes.
36 Mistral Large 3 68 Mistral AI 256K tokens $0.50 $1.50 European/open-license deployments and value-focused production apps.
37 DeepSeek V4 Flash 66 DeepSeek 1M tokens (384K max output) $0.14 ($0.0028 cache hit) $0.28 High-volume, repetitive workloads and budget pipelines needing long context.
38 Doubao Seed 2.1 Turbo New 65 ByteDance ≈256K tokens $0.45 $2.25 High-volume, cost-sensitive coding and agent workloads needing near-flagship quality at low cost.
39 Grok 4.1 Fast 64 xAI 2M tokens $0.20 ($0.05 cached) $0.50 Large-context, high-volume agentic workloads where cost and context size dominate.
40 Gemini 3 Flash 62 Google 1M tokens $0.50 $3.00 Cost-sensitive multimodal and long-context tasks at scale.
41 GPT-5.4 mini 60 OpenAI ~400K tokens $0.75 $4.50 High-volume tasks like classification, extraction, and chat where cost and speed matter most.
42 Claude Haiku 4.5 58 Anthropic 200K tokens $1.00 $5.00 Real-time chat, routing, classification, and high-throughput pipelines on a budget.
43 Amazon Nova 2 Pro New 56 Amazon 256K tokens $1.25 $10.00 AWS-centric enterprises wanting a frontier-class model inside Bedrock.
44 Llama 4 Scout 54 Meta 10M tokens ≈$0.08 ≈$0.30 Extreme long-context tasks (whole codebases, large document sets) on an open, cheap model.
45 Gemini 2.5 Flash-Lite New 52 Google 1M tokens (~64K output) $0.10 $0.40 High-volume classification, summarization, extraction and simple chat at scale (note the Oct 2026 Gemini API retirement).
46 Amazon Nova Premier Updated 45 Amazon 1M tokens $2.50 $12.50 AWS-native enterprises needing a managed multimodal model with strong governance.

Ranked by the Benchquill score — a 0–100 capability-first editorial index weighing reasoning quality, agentic & coding strength, context window and price-performance (how we rank). Prices are public API list prices per 1M tokens (USD) and may change — verify with each provider. July 2, 2026.

Not sure which model to pick?

Compare any two models side-by-side on price, context and capability.

Compare models