LLM Landscape 2026: Intelligence Leaderboard and Model Guide
Leaderboard Methodology
Both tables use the AA (Artificial Analysis) Intelligence Index v4.1, which aggregates nine agentic-weighted evaluations — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR — into a single normalized integer. v4.1 re-baselines scores versus v4.0: a model that scored 61 under v4.0 may land near 56 under v4.1; compare within a methodology version, not across them. Scores shown use each model's highest published effort tier (typically max / xhigh). Context windows use comma-separated token counts; missing public data is "—". Pricing is per million tokens (input / output).
Table 1 — Top 25 by vendor: exactly one representative per company — the provider's newest clearly superior flagship or best overall product — to surface geographic and corporate diversity. Table 2 — Top 25 by model: pure AA Index ranking; multiple entries from Anthropic, OpenAI, Google, and others are expected. Models with limited or suspended API access (Fable 5, Mythos 5, GPT-5.6 Sol preview) are included with availability notes because they materially affect the competitive picture.
Top 25 LLMs by Vendor — Company Diversity Leaderboard (July 2026)
| Rank | Model | Capability Index (AA Index) | Context Window (tokens) | Input Cost ($/M tokens) | Output Cost ($/M tokens) | Notes |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.8 Anthropic | 56 | 1,000,000 | $5.00 | $25.00 | Best generally available Anthropic model; Sonnet 5 (AA 53) now default for Free/Pro; Fable/Mythos 5 (AA 60) suspended with staged return |
| 2 | GPT-5.5 OpenAI | 55 | 1,050,000 | $5.00 | $30.00 | Flagship API model (xhigh); GPT-5.6 Sol preview (June 26) gated to trusted partners pending broader GA |
| 3 | GLM-5.2 Z.ai | 51 | 1,000,000 | $1.40 | $4.40 | Open Weight Leading open-weight AA model; SWE-bench Pro 62.1%, Terminal-Bench 2.1 81.0%; 744B MoE (40B active) |
| 4 | Gemini 3.5 Flash Google | 50 | 1,000,000 | $1.50 | $9.00 | Google API flagship; 280+ t/s; Gemma 4 12B/31B cover open-weight (see model table). Gemini 3.5 Pro delayed to July |
| 5 | DeepSeek V4 Pro DeepSeek AI | 44 | 1,000,000 | $2.19 | $8.76 | Open Weight 1.6T MoE (49B active); 1M context; V4 Flash for cost-sensitive pipelines |
| 6 | MiniMax-M3 MiniMax | 44 | 1,000,000 | — | — | Open Weight Tied DeepSeek V4 Pro on AA v4.1; strong SWE-bench Verified (~80.5%) |
| 7 | Kimi K2.6 Moonshot AI | 43 | — | — | — | Open Weight 1T params; Kimi K2.7-Code (June 12) targets long-horizon coding |
| 8 | MiMo-V2.5-Pro Xiaomi | 42 | 1,000,000 | — | — | Successor to MiMo-V2-Pro; pricing not publicly disclosed |
| 9 | Grok 4.3 xAI | 38 | 1,000,000 | $1.25 | $2.50 | Fastest AA task completion (~1.5 min); Grok 4.20 still used for multi-agent real-time loops |
| 10 | Qwen3.5 397B A17B Alibaba | 34 | 262,000 | $0.60 | $3.60 | Open Weight Apache 2.0; best Qwen-family AA representative |
| 11 | NVIDIA Nemotron 3 Super 120B NVIDIA | 25 | 1,000,000 | $0.30 | $0.75 | Open Weight Enterprise self-hosting value tier; Mamba-2/MoE architecture |
| 12 | Mistral Large 3 Mistral | 16 | 256,000 | $0.50 | $1.50 | Open Weight EU-based; Apache 2.0; v4.1 agentic re-weighting lowered composite vs v4.0 |
| 13 | Nova Premier Amazon | 13 | 1,000,000 | $2.50 | $12.50 | Hyperscaler representative; deep AWS integration |
| 14 | Llama 4 Scout Meta | 10 | 10,000,000 | — | — | Open Weight Context-window outlier; 10M tokens for corpus-scale self-hosting |
| 15 | Command A Cohere | 8 | 256,000 | $2.50 | $10.00 | Enterprise RAG and tool-use focus |
| 16 | Solar Pro 2 Upstage | 8 | — | — | — | 31B Korean frontier; strong regional/language performance; v4.1 composite lower than legacy v4.0 ranking |
| 17 | ERNIE 4.5 300B A47B Baidu | — | — | — | — | Best verifiable ERNIE-family public entry; AA v4.1 pending |
| 18 | Granite 4.0 H Small IBM | 5 | — | — | — | Open Weight Enterprise governance and open-deployment focus |
| 19 | Jamba 1.7 Large AI21 | 5 | — | — | — | Hybrid SSM/Transformer architecture for long-input efficiency |
| 20 | Yi-Lightning 01.AI | — | — | — | — | Vendor-diversity slot; public AA v4.1 score not yet published |
| 21 | Sonar Reasoning Pro Perplexity | — | 128,000 | — | — | Search-augmented reasoning API; AA v4.1 pending |
| 22 | Reka Flash 3 Reka | — | — | — | — | Multimodal agentic model; AA v4.1 pending |
| 23 | Hunyuan-A13B-Instruct Tencent | — | — | — | — | Chinese hyperscaler representative; AA v4.1 pending |
| 24 | Stable LM 2 12B Stability AI | — | — | — | — | Open Weight Community/open-deployment slot; AA v4.1 pending |
| 25 | Evo-Ukiyoe Sakana AI | — | — | — | — | Evolutionary-model research lab; specialized rather than general-purpose frontier |
Top 25 Models by AA Index v4.1 — Pure Capability Leaderboard (July 2026)
This table ranks the twenty-five highest-scoring models on the AA Index v4.1 regardless of vendor — expect multiple Anthropic, OpenAI, and Google entries. Effort tiers are max / xhigh unless noted.
| Rank | Model | Capability Index (AA v4.1) | Context Window (tokens) | Input Cost ($/M tokens) | Output Cost ($/M tokens) | Notes |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 Anthropic | 60 | 1,000,000 | $10.00 | $50.00 | Limited access Mythos-class with safety classifiers; suspended June 12, staged return expected; Opus 4.8 fallback on blocked queries |
| 2 | Claude Opus 4.8 Anthropic | 56 | 1,000,000 | $5.00 | $25.00 | Best generally available model; SWE-bench Pro 69.2%; leads available tier |
| 3 | GPT-5.5 (xhigh) OpenAI | 55 | 1,050,000 | $5.00 | $30.00 | OpenAI API flagship; Terminal-Bench 2.1 leader among closed models |
| 4 | Claude Opus 4.7 Anthropic | 54 | 1,000,000 | $5.00 | $25.00 | Prior Opus generation; still strong for pinned production workflows |
| 5 | Claude Sonnet 5 Anthropic | 53 | 1,000,000 | $2.00 | $10.00 | New Jul 1 Default Free/Pro model; Terminal-Bench 2.1 80.4% beats Opus 4.8; intro pricing through Aug 31 |
| 6 | GPT-5.5 (high) OpenAI | 53 | 922,000 | $5.00 | $30.00 | Same family as xhigh at lower reasoning depth; more token-efficient on AA tasks |
| 7 | GLM-5.2 (max) Z.ai | 51 | 1,000,000 | $1.40 | $4.40 | Open Weight #1 open-weight AA; SWE-bench Pro 62.1%; MIT license; June 13 GA |
| 8 | GPT-5.4 OpenAI | 51 | 1,050,000 | $2.50 | $15.00 | Prior flagship; still viable at half the output cost of GPT-5.5 |
| 9 | Gemini 3.5 Flash (high) Google | 50 | 1,000,000 | $1.50 | $9.00 | Speed-intelligence Pareto leader; 163+ t/s; +9 pts vs Gemini 3 Flash on v4.1 |
| 10 | Claude Sonnet 4.6 (max) Anthropic | 47 | 1,000,000 | $3.00 | $15.00 | Superseded by Sonnet 5; still pinned in production agent stacks |
| 11 | Gemini 3.1 Pro Preview Google | 46 | 1,000,000 | $2.00 | $12.00 | Fastest per-task completion (~1.6 min) among top-tier models; multimodal strength |
| 12 | Gemini 3.5 Flash (medium) Google | 45 | 1,000,000 | $1.50 | $9.00 | Default Flash effort tier; balances speed and cost for high-volume agents |
| 13 | DeepSeek V4 Pro (max) DeepSeek AI | 44 | 1,000,000 | $2.19 | $8.76 | Open Weight Tied MiniMax-M3; best cost-per-task among open weights on AA v4.1 |
| 14 | MiniMax-M3 MiniMax | 44 | 1,000,000 | — | — | Open Weight Tied DeepSeek V4 Pro; strong multimodal and agentic scores |
| 15 | Kimi K2.6 Moonshot AI | 43 | — | — | — | Open Weight 1T params; K2.7-Code adds +21.8% on Kimi Code Bench v2 |
| 16 | MiMo-V2.5-Pro Xiaomi | 42 | 1,000,000 | — | — | 1M context; top-tier Chinese proprietary entrant |
| 17 | Grok 4.3 (high) xAI | 38 | 1,000,000 | $1.25 | $2.50 | Fastest wall-clock per AA task (~1.5 min); excellent $/capability ratio |
| 18 | Grok 4.20 xAI | 37 | 2,000,000 | $2.00 | $6.00 | 2M context; multi-agent real-time workflows |
| 19 | Qwen3.5 397B A17B Alibaba | 34 | 262,000 | $0.60 | $3.60 | Open Weight Apache 2.0; 119-language support |
| 20 | o3 OpenAI | 30 | 200,000 | $2.00 | $8.00 | Dedicated reasoning line; strong GPQA/AIME at mid-tier pricing |
| 21 | Gemma 4 31B Google (open-weight) | 29 | 256,000 | — | — | Open Weight Top Gemma 4 composite; Apache 2.0; multimodal text/image/video |
| 22 | o4-mini OpenAI | 26 | 200,000 | $1.10 | $4.40 | Cost-efficient reasoning; 142 t/s throughput |
| 23 | NVIDIA Nemotron 3 Super 120B NVIDIA | 25 | 1,000,000 | $0.30 | $0.75 | Open Weight Best $/M among 1M-context open models |
| 24 | gpt-oss-120B OpenAI (open-weight) | 24 | 128,000 | $0.30 | $0.30 | Open Weight Managed API open-weight access at $0.30/M |
| 25 | Gemma 4 12B Google (open-weight) | 22 | 256,000 | — | — | Open Weight Encoder-free multimodal; 16GB laptop deploy; MMLU-Pro 77.2%; Apache 2.0 (June 3) |
Claude Mythos 5 shares Fable 5's underlying weights and AA score (~60) but is restricted to Project Glasswing partners (cyber/biology safeguards lifted). GPT-5.6 Sol preview (June 26) is gated to government-vetted partners; Terra and Luna tiers offer lower cost. Neither appears in the AA v4.1 table yet. Gemma 4 12B is a deployment play — 77.2% MMLU-Pro at ~6.6 GB VRAM (Q4) — not an AA leaderboard climber, but the most important open-weight laptop release of the quarter.
Key Takeaways
Key Performance Metrics
Task-Specific Leaders
| Model | Benchmark Leadership |
|---|---|
| Claude Fable 5 | AA v4.1 peak (60) · limited access · Mythos-class |
| Claude Opus 4.8 | Best available · SWE-Bench Pro 69.2% · AA 56 |
| Claude Sonnet 5 | Terminal-Bench 2.1 80.4% · AA 53 · new default |
| GLM-5.2 | Open-weight AA leader · SWE-bench Pro 62.1% · AA 51 |
| GPT-5.5 (xhigh) | Terminal-Bench 2.1 · closed-model leader · AA 55 |
| Gemini 3.5 Flash | Speed-intelligence Pareto · 163+ t/s · AA 50 |
Context Window Champions
| Model | Tokens |
|---|---|
| Llama 4 Scout | 10,000,000 |
| GLM-5.2 | 1,000,000 |
| GPT-5.5 · Gemini 3.5 · Claude · MiMo · DeepSeek V4 Pro · Nemotron | 1,000,000–1,050,000 |
| Gemma 4 12B / 31B | 256,000 |
Cost Efficiency
| Tier | Models | Output $/M |
|---|---|---|
| Best Value | Nemotron · DeepSeek V4 Pro · GLM-5.2 | ~$0.75–$4.40 |
| Mid-Range | Sonnet 5 · Gemini 3.5 Flash · Qwen3.5 | $3.60–$10.00 |
| Flagship | Gemini 3.1 Pro · Grok 4.3 | $2.50–$12.00 |
| Premium | GPT-5.5 · Claude Opus 4.8 · Fable 5 | $25.00–$50.00 |
Specialized Performance Highlights
Speed & Latency
Open-Weight Excellence
| Model | Key Strength |
|---|---|
| GLM-5.2 | AA 51 · SWE-bench Pro 62.1% · 1M ctx · MIT |
| Gemma 4 12B | Laptop-class · encoder-free multimodal · Apache 2.0 · AA 22 |
| Gemma 4 31B | AA 29 · multimodal · 256K ctx |
| DeepSeek V4 Pro | AA 44 · 1M context · best $/task open weight |
| MiniMax-M3 | AA 44 · tied DeepSeek · multimodal agents |
| Llama 4 Scout | 10M-token context · corpus-scale tasks |
| Kimi K2.6 / K2.7-Code | AA 43 · long-horizon coding specialist |
Model Selection Guide
Industry Impact & Future Trends (2026)
The 2026 LLM landscape is defined by AA Index v4.1's agentic re-weighting, regulatory friction on frontier releases, and a collapsing mid-tier:
Coding & Agents
Regulatory & Access
Open-Weight & Local AI
Conclusion
The July 2026 landscape shifted in four weeks. AA Index v4.1 re-baselined scores with heavier agentic weighting — compare within v4.1, not against June's v4.0 numbers. Anthropic launched Sonnet 5 as the default mid-tier agent, while Fable/Mythos 5 briefly topped the index at AA 60 before a government-mandated suspension. Z.ai's GLM-5.2 became the leading open-weight model (AA 51, SWE-bench Pro 62.1%). Google's Gemma 4 12B redefined laptop-class multimodal deployment. OpenAI previewed GPT-5.6 Sol under restricted access; Gemini 3.5 Pro is delayed to July.
Strategic Takeaway (2026)
Looking ahead: Gemini 3.5 Pro and GPT-5.6 general availability will reshuffle both tables. AA v4.1's agentic focus means Terminal-Bench and GDPval-AA matter more than static knowledge benchmarks. Expect another refresh when those models ship broadly — likely within weeks of this update.