Leaderboard

AI model leaderboard 2026

47 leading large language models compared on context window, API price and best-fit use case. Updated July 30, 2026.

Quick answer

As of July 30, 2026, the leaderboard spans 47 models from 20 developers. Context windows reach 10M tokens; API prices range widely — see the table, or compare any two.

#ModelBq scoreDeveloperContextInput /1MOutput /1MBest for
01 Claude Fable 5 New 98 Anthropic 1M tokens (128K max output) $10.00 $50.00 The hardest multi-step agentic and coding work, frontier reasoning, and long-running autonomous tasks where capability matters most.
02 Claude Opus 5 New 97 Anthropic 1M tokens (128K max output) $5.00 $25.00 Frontier-grade reasoning, deep agents and hard coding work where you want Fable-5-class quality without Fable-5 pricing.
03 GPT-5.6 Sol New 96 OpenAI ~1M tokens (128K max output) $5.00 $30.00 The highest-stakes reasoning, deep agents and frontier coding where capability outweighs cost.
04 GPT-5.5 95 OpenAI ~1.05M tokens (128K max output) $5.00 $30.00 Highest-stakes reasoning, complex agents, and frontier coding tasks where capability outweighs cost.
05 Claude Opus 4.8 Updated 93 Anthropic 1M tokens $5.00 $25.00 Complex coding agents, long-horizon tasks, and workloads needing top reliability and reasoning.
06 Gemini 3.1 Pro 92 Google 2M tokens $2.00 (under 200K; $4.00 above) $12.00 (under 200K; $18.00 above) Long-document and multimodal workloads, RAG over huge corpora, and Google Cloud-native apps.
07 Grok 4.5 New 91 xAI 500K tokens $2.00 $6.00 Coding agents and long-running tool-use workflows that want a frontier model at mid-tier pricing.
08 Claude Opus 4.7 New 90 Anthropic 1M tokens $5.00 $25.00 Production coding agents, long-horizon workflow automation, and document or finance agent work.
09 GPT-5.6 Terra New 90 OpenAI ~1M tokens (128K max output) $2.50 $15.00 Most production apps that want near-flagship quality at a materially lower per-token cost.
10 Grok 4.3 88 xAI 1M tokens $1.25 ($0.20 cached) $2.50 Reasoning and agentic tasks needing current/real-time info and low output costs.
11 Claude Sonnet 5 New 87 Anthropic 1M tokens $2.00 $10.00 Running AI agents affordably, multi-step automation, coding assistants and high-scale chat where you want near-flagship quality at lower cost.
12 Kimi K3 New 87 Moonshot AI 1,048,576 tokens $3.00 $15.00 Teams that want frontier-adjacent, multimodal capability with the option to self-host rather than stay on a closed API.
13 GPT-5.4 86 OpenAI ~1M tokens (922K input, 128K output) $2.50 $15.00 Most production apps wanting near-flagship quality at roughly half the flagship price.
14 Meta Muse Spark 1.1 New 85 Meta 1,048,576 tokens $1.25 $4.25 Cost-sensitive coding and computer-use agents that need long context and a drop-in OpenAI-compatible API.
15 DeepSeek V4 Pro 84 DeepSeek 1M tokens (384K max output) $0.435 $0.87 Cost-conscious teams wanting frontier reasoning and coding at a fraction of US-lab prices.
16 MiMo-V2.5-Pro New 84 Xiaomi 1M tokens $0.435 $0.87 Agentic coding and long-running autonomous agents on a budget, including self-hosted deployments.
17 GPT-5.6 Luna New 83 OpenAI ~1M tokens (128K max output) $1.00 $6.00 High-volume, latency-sensitive workloads and cheap agent loops where flagship depth isn't required.
18 MiMo-V2.5 New 83 Xiaomi 1M tokens $0.105 $0.28 High-volume agentic pipelines, long-document and multimodal processing where cost matters most.
19 Tencent Hy3 New 83 Tencent (Hunyuan) 256K tokens $0.063 $0.21 Cost-efficient agent pipelines, coding agents and high-volume production workloads.
20 Qwen3.7 Max Updated 82 Alibaba 1M tokens (64K max output) $1.25 $3.75 Long-context agentic workloads, multilingual apps, and Asia-Pacific deployments.
21 Gemini 3.6 Flash New 81 Google 1M tokens $1.50 $7.50 High-volume production work and agentic workflows that need solid coding and document skills at Flash-tier cost.
22 Kimi K2.7 Code New 81 Moonshot AI 256K tokens $0.95 $4.00 Repository-scale refactoring and long multi-turn agentic coding with an open-weight, cost-efficient model.
23 Z.ai GLM-5.2 New 80 Z.ai (Zhipu AI) 1M tokens $0.95 $3.00 Cost-efficient agentic coding and self-hosted deployments where open weights matter.
24 MiniMax M3 New 79 MiniMax 1M tokens (min 512K) $0.60 $2.40 Long-context coding agents, computer-use automation and high-volume tasks wanting frontier quality with open weights.
25 Doubao Seed 2.1 Pro New 78 ByteDance Not disclosed $0.83 $4.17 Complex coding and long-chain agent tasks at low cost, especially for teams already on Volcano Engine.
26 Gemini 3.5 Flash Updated 77 Google 1M tokens $1.50 $9.00 High-volume multimodal apps and agents needing speed with large context.
27 Sakana Fugu Ultra New 76 Sakana AI 1M tokens (128K max output; tier changes above 272K) $5.00 $30.00 Hard multi-step reasoning, coding and code review, scientific/research tasks, and quality-critical agentic workflows.
28 Claude Sonnet 4.6 Updated 75 Anthropic 1M tokens $3.00 $15.00 Everyday coding, agents, and production apps wanting Claude quality without flagship pricing.
29 Inkling New 73 Thinking Machines Lab Up to 1M tokens ≈$1.00 ≈$4.05 Research teams and companies that need to fine-tune and fully control a large multimodal model on their own infrastructure.
30 Laguna M.1 New 73 Poolside 262K tokens (32K max output) $0.20 $0.40 Agentic software engineering: codebase exploration, multi-file edits, test loops and CLI agents.
31 Kimi K2.6 Updated 72 Moonshot AI 262K tokens (64K max output) ≈$0.60 ≈$2.50 Open-model coding agents and budget-conscious agentic pipelines.
32 Llama 4 Maverick 70 Meta 1M tokens ≈$0.15 ≈$0.60 Teams wanting an open, multimodal model they can self-host or run cheaply via multiple providers.
33 Nemotron 3 Ultra New 70 NVIDIA 1M tokens $0.50 $2.20 Long-running agent workflows, long-context processing and cost-efficient self-hosted deployments.
34 Step 3.7 Flash New 70 StepFun 256K tokens $0.20 $1.15 High-volume, cost-sensitive agents, search workflows and structured outputs where speed matters most.
35 gpt-oss-120b New 69 OpenAI 131K tokens ≈$0.03 ≈$0.15 Cost-sensitive high-volume workloads, on-prem or privacy-constrained deployments, and custom fine-tunes.
36 Mistral Large 3 68 Mistral AI 256K tokens $0.50 $1.50 European/open-license deployments and value-focused production apps.
37 DeepSeek V4 Flash 66 DeepSeek 1M tokens (384K max output) $0.14 ($0.0028 cache hit) $0.28 High-volume, repetitive workloads and budget pipelines needing long context.
38 Doubao Seed 2.1 Turbo New 65 ByteDance ≈256K tokens $0.45 $2.25 High-volume, cost-sensitive coding and agent workloads needing near-flagship quality at low cost.
39 Grok 4.1 Fast 64 xAI 2M tokens $0.20 ($0.05 cached) $0.50 Large-context, high-volume agentic workloads where cost and context size dominate.
40 Gemini 3.5 Flash-Lite New 63 Google 1M tokens $0.30 $2.50 Bulk classification, extraction and document pipelines where throughput and unit cost matter more than depth.
41 Gemini 3 Flash 62 Google 1M tokens $0.50 $3.00 Cost-sensitive multimodal and long-context tasks at scale.
42 GPT-5.4 mini 60 OpenAI ~400K tokens $0.75 $4.50 High-volume tasks like classification, extraction, and chat where cost and speed matter most.
43 Claude Haiku 4.5 58 Anthropic 200K tokens $1.00 $5.00 Real-time chat, routing, classification, and high-throughput pipelines on a budget.
44 Amazon Nova 2 Pro New 56 Amazon 256K tokens $1.25 $10.00 AWS-centric enterprises wanting a frontier-class model inside Bedrock.
45 Llama 4 Scout 54 Meta 10M tokens ≈$0.08 ≈$0.30 Extreme long-context tasks (whole codebases, large document sets) on an open, cheap model.
46 Gemini 2.5 Flash-Lite New 52 Google 1M tokens (~64K output) $0.10 $0.40 High-volume classification, summarization, extraction and simple chat at scale (note the Oct 2026 Gemini API retirement).
47 Amazon Nova Premier Updated 45 Amazon 1M tokens $2.50 $12.50 AWS-native enterprises needing a managed multimodal model with strong governance.

Ranked by the Benchquill score — a 0–100 capability-first editorial index weighing reasoning quality, agentic & coding strength, context window and price-performance (how we rank). Prices are public API list prices per 1M tokens (USD) and may change — verify with each provider. July 30, 2026.

Not sure which model to pick?

Compare any two models side-by-side on price, context and capability.

Compare models