LLM Landscape 2026: Intelligence Leaderboard and Model Guide

A July 2026 snapshot of the frontier LLM landscape, now with two Top 25 leaderboards: one model per vendor (company diversity) and a pure model ranking by AA Intelligence Index v4.1 score (multiple models per lab allowed). Since June, Anthropic shipped Claude Sonnet 5 (AA 53, June 30) as the new default mid-tier agent; Z.ai's GLM-5.2 (AA 51) became the leading open-weight model on SWE-bench Pro and the AA composite; Google's Gemma 4 12B (June 3) brought encoder-free multimodal intelligence to laptop-class hardware under Apache 2.0; and Anthropic's Mythos/Fable 5 pair briefly topped the index at AA 60 before a government-mandated suspension — access is returning in stages. OpenAI previewed GPT-5.6 Sol (June 26) under restricted partner access; Gemini 3.5 Pro is delayed to July.

Leaderboard Methodology

Both tables use the AA (Artificial Analysis) Intelligence Index v4.1, which aggregates nine agentic-weighted evaluations — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR — into a single normalized integer. v4.1 re-baselines scores versus v4.0: a model that scored 61 under v4.0 may land near 56 under v4.1; compare within a methodology version, not across them. Scores shown use each model's highest published effort tier (typically max / xhigh). Context windows use comma-separated token counts; missing public data is "—". Pricing is per million tokens (input / output).

Table 1 — Top 25 by vendor: exactly one representative per company — the provider's newest clearly superior flagship or best overall product — to surface geographic and corporate diversity. Table 2 — Top 25 by model: pure AA Index ranking; multiple entries from Anthropic, OpenAI, Google, and others are expected. Models with limited or suspended API access (Fable 5, Mythos 5, GPT-5.6 Sol preview) are included with availability notes because they materially affect the competitive picture.

Top 25 LLMs by Vendor — Company Diversity Leaderboard (July 2026)

RankModelCapability Index
(AA Index)
Context Window
(tokens)
Input Cost
($/M tokens)
Output Cost
($/M tokens)
Notes
1
Claude Opus 4.8
Anthropic
561,000,000$5.00$25.00Best generally available Anthropic model; Sonnet 5 (AA 53) now default for Free/Pro; Fable/Mythos 5 (AA 60) suspended with staged return
2
GPT-5.5
OpenAI
551,050,000$5.00$30.00Flagship API model (xhigh); GPT-5.6 Sol preview (June 26) gated to trusted partners pending broader GA
3
GLM-5.2
Z.ai
511,000,000$1.40$4.40Open Weight Leading open-weight AA model; SWE-bench Pro 62.1%, Terminal-Bench 2.1 81.0%; 744B MoE (40B active)
4
Gemini 3.5 Flash
Google
501,000,000$1.50$9.00Google API flagship; 280+ t/s; Gemma 4 12B/31B cover open-weight (see model table). Gemini 3.5 Pro delayed to July
5
DeepSeek V4 Pro
DeepSeek AI
441,000,000$2.19$8.76Open Weight 1.6T MoE (49B active); 1M context; V4 Flash for cost-sensitive pipelines
6
MiniMax-M3
MiniMax
441,000,000Open Weight Tied DeepSeek V4 Pro on AA v4.1; strong SWE-bench Verified (~80.5%)
7
Kimi K2.6
Moonshot AI
43Open Weight 1T params; Kimi K2.7-Code (June 12) targets long-horizon coding
8
MiMo-V2.5-Pro
Xiaomi
421,000,000Successor to MiMo-V2-Pro; pricing not publicly disclosed
9
Grok 4.3
xAI
381,000,000$1.25$2.50Fastest AA task completion (~1.5 min); Grok 4.20 still used for multi-agent real-time loops
10
Qwen3.5 397B A17B
Alibaba
34262,000$0.60$3.60Open Weight Apache 2.0; best Qwen-family AA representative
11
NVIDIA Nemotron 3 Super 120B
NVIDIA
251,000,000$0.30$0.75Open Weight Enterprise self-hosting value tier; Mamba-2/MoE architecture
12
Mistral Large 3
Mistral
16256,000$0.50$1.50Open Weight EU-based; Apache 2.0; v4.1 agentic re-weighting lowered composite vs v4.0
13
Nova Premier
Amazon
131,000,000$2.50$12.50Hyperscaler representative; deep AWS integration
14
Llama 4 Scout
Meta
1010,000,000Open Weight Context-window outlier; 10M tokens for corpus-scale self-hosting
15
Command A
Cohere
8256,000$2.50$10.00Enterprise RAG and tool-use focus
16
Solar Pro 2
Upstage
831B Korean frontier; strong regional/language performance; v4.1 composite lower than legacy v4.0 ranking
17
ERNIE 4.5 300B A47B
Baidu
Best verifiable ERNIE-family public entry; AA v4.1 pending
18
Granite 4.0 H Small
IBM
5Open Weight Enterprise governance and open-deployment focus
19
Jamba 1.7 Large
AI21
5Hybrid SSM/Transformer architecture for long-input efficiency
20
Yi-Lightning
01.AI
Vendor-diversity slot; public AA v4.1 score not yet published
21
Sonar Reasoning Pro
Perplexity
128,000Search-augmented reasoning API; AA v4.1 pending
22
Reka Flash 3
Reka
Multimodal agentic model; AA v4.1 pending
23
Hunyuan-A13B-Instruct
Tencent
Chinese hyperscaler representative; AA v4.1 pending
24
Stable LM 2 12B
Stability AI
Open Weight Community/open-deployment slot; AA v4.1 pending
25
Evo-Ukiyoe
Sakana AI
Evolutionary-model research lab; specialized rather than general-purpose frontier

Top 25 Models by AA Index v4.1 — Pure Capability Leaderboard (July 2026)

This table ranks the twenty-five highest-scoring models on the AA Index v4.1 regardless of vendor — expect multiple Anthropic, OpenAI, and Google entries. Effort tiers are max / xhigh unless noted.

RankModelCapability Index
(AA v4.1)
Context Window
(tokens)
Input Cost
($/M tokens)
Output Cost
($/M tokens)
Notes
1
Claude Fable 5
Anthropic
601,000,000$10.00$50.00Limited access Mythos-class with safety classifiers; suspended June 12, staged return expected; Opus 4.8 fallback on blocked queries
2
Claude Opus 4.8
Anthropic
561,000,000$5.00$25.00Best generally available model; SWE-bench Pro 69.2%; leads available tier
3
GPT-5.5 (xhigh)
OpenAI
551,050,000$5.00$30.00OpenAI API flagship; Terminal-Bench 2.1 leader among closed models
4
Claude Opus 4.7
Anthropic
541,000,000$5.00$25.00Prior Opus generation; still strong for pinned production workflows
5
Claude Sonnet 5
Anthropic
531,000,000$2.00$10.00New Jul 1 Default Free/Pro model; Terminal-Bench 2.1 80.4% beats Opus 4.8; intro pricing through Aug 31
6
GPT-5.5 (high)
OpenAI
53922,000$5.00$30.00Same family as xhigh at lower reasoning depth; more token-efficient on AA tasks
7
GLM-5.2 (max)
Z.ai
511,000,000$1.40$4.40Open Weight #1 open-weight AA; SWE-bench Pro 62.1%; MIT license; June 13 GA
8
GPT-5.4
OpenAI
511,050,000$2.50$15.00Prior flagship; still viable at half the output cost of GPT-5.5
9
Gemini 3.5 Flash (high)
Google
501,000,000$1.50$9.00Speed-intelligence Pareto leader; 163+ t/s; +9 pts vs Gemini 3 Flash on v4.1
10
Claude Sonnet 4.6 (max)
Anthropic
471,000,000$3.00$15.00Superseded by Sonnet 5; still pinned in production agent stacks
11
Gemini 3.1 Pro Preview
Google
461,000,000$2.00$12.00Fastest per-task completion (~1.6 min) among top-tier models; multimodal strength
12
Gemini 3.5 Flash (medium)
Google
451,000,000$1.50$9.00Default Flash effort tier; balances speed and cost for high-volume agents
13
DeepSeek V4 Pro (max)
DeepSeek AI
441,000,000$2.19$8.76Open Weight Tied MiniMax-M3; best cost-per-task among open weights on AA v4.1
14
MiniMax-M3
MiniMax
441,000,000Open Weight Tied DeepSeek V4 Pro; strong multimodal and agentic scores
15
Kimi K2.6
Moonshot AI
43Open Weight 1T params; K2.7-Code adds +21.8% on Kimi Code Bench v2
16
MiMo-V2.5-Pro
Xiaomi
421,000,0001M context; top-tier Chinese proprietary entrant
17
Grok 4.3 (high)
xAI
381,000,000$1.25$2.50Fastest wall-clock per AA task (~1.5 min); excellent $/capability ratio
18
Grok 4.20
xAI
372,000,000$2.00$6.002M context; multi-agent real-time workflows
19
Qwen3.5 397B A17B
Alibaba
34262,000$0.60$3.60Open Weight Apache 2.0; 119-language support
20
o3
OpenAI
30200,000$2.00$8.00Dedicated reasoning line; strong GPQA/AIME at mid-tier pricing
21
Gemma 4 31B
Google (open-weight)
29256,000Open Weight Top Gemma 4 composite; Apache 2.0; multimodal text/image/video
22
o4-mini
OpenAI
26200,000$1.10$4.40Cost-efficient reasoning; 142 t/s throughput
23
NVIDIA Nemotron 3 Super 120B
NVIDIA
251,000,000$0.30$0.75Open Weight Best $/M among 1M-context open models
24
gpt-oss-120B
OpenAI (open-weight)
24128,000$0.30$0.30Open Weight Managed API open-weight access at $0.30/M
25
Gemma 4 12B
Google (open-weight)
22256,000Open Weight Encoder-free multimodal; 16GB laptop deploy; MMLU-Pro 77.2%; Apache 2.0 (June 3)

Claude Mythos 5 shares Fable 5's underlying weights and AA score (~60) but is restricted to Project Glasswing partners (cyber/biology safeguards lifted). GPT-5.6 Sol preview (June 26) is gated to government-vetted partners; Terra and Luna tiers offer lower cost. Neither appears in the AA v4.1 table yet. Gemma 4 12B is a deployment play — 77.2% MMLU-Pro at ~6.6 GB VRAM (Q4) — not an AA leaderboard climber, but the most important open-weight laptop release of the quarter.

Key Takeaways

Peak Intelligence
Claude Opus 4.8 leads the available tier at AA 56; Fable 5 briefly held AA 60 before suspension. GPT-5.5 xhigh (55) and Sonnet 5 (53, launched June 30) compress the mid-tier — Sonnet 5 beats Opus 4.8 on Terminal-Bench 2.1 (80.4% vs 74.6%). Gemini 3.5 Pro delayed to July; GPT-5.6 Sol preview gated to partners.
Coding & Agentic Leadership
GLM-5.2 leads open-weight SWE-bench Pro (62.1%) and AA composite (51). Claude Opus 4.8 still leads closed-model SWE-bench Pro (69.2%). Sonnet 5 is the new default for agentic coding at half Opus cost. Grok 4.3 completes AA tasks in ~1.5 minutes — fastest wall-clock in the top tier.
Context-Window Outlier
Llama 4 Scout pushes open-weight context to 10,000,000 tokens — enabling full-codebase and corpus-scale analysis in a single pass. The top closed flagships cluster at 1,000,000–1,050,000 tokens.
Cost-Efficient Frontier
DeepSeek V4 Pro delivers AA 44 at $0.04/task on AA v4.1 — best open-weight cost efficiency. Nemotron 3 Super ($0.30 / $0.75) anchors API self-hosting value. Sonnet 5 intro pricing ($2 / $10 through Aug 31) makes all-day agents economically viable. Gemini 3.5 Flash ($1.50 / $9) remains the speed/value sweet spot.
Emerging Challengers
GLM-5.2 (AA 51) and Sonnet 5 (AA 53) reshuffled the mid-tier. Kimi K2.7-Code targets long-horizon coding. GPT-5.6 Sol/Terra/Luna preview (June 26) awaits broader release. Mythos/Fable 5 access returning after June suspension. Gemini 3.5 Pro expected July.
Open-Weight Size Efficiency (Gemma 4)
Gemma 4 12B (June 3, Apache 2.0) is the headline open-weight release: encoder-free multimodal (text/image/audio), 256K context, MMLU-Pro 77.2%, runs on 16GB laptops. AA 22 — a deployment play, not a composite leader. Gemma 4 31B (AA 29) remains the family's AA benchmark. GLM-5.2 displaced Kimi K2.6 as the #1 open-weight AA model.

Key Performance Metrics

Task-Specific Leaders
ModelBenchmark Leadership
Claude Fable 5AA v4.1 peak (60) · limited access · Mythos-class
Claude Opus 4.8Best available · SWE-Bench Pro 69.2% · AA 56
Claude Sonnet 5Terminal-Bench 2.1 80.4% · AA 53 · new default
GLM-5.2Open-weight AA leader · SWE-bench Pro 62.1% · AA 51
GPT-5.5 (xhigh)Terminal-Bench 2.1 · closed-model leader · AA 55
Gemini 3.5 FlashSpeed-intelligence Pareto · 163+ t/s · AA 50
Context Window Champions
ModelTokens
Llama 4 Scout10,000,000
GLM-5.21,000,000
GPT-5.5 · Gemini 3.5 · Claude · MiMo · DeepSeek V4 Pro · Nemotron1,000,000–1,050,000
Gemma 4 12B / 31B256,000
10M tokens fits entire codebases; 1M+ handles legal corpora and research archives in one pass
Cost Efficiency
TierModelsOutput $/M
Best ValueNemotron · DeepSeek V4 Pro · GLM-5.2~$0.75–$4.40
Mid-RangeSonnet 5 · Gemini 3.5 Flash · Qwen3.5$3.60–$10.00
FlagshipGemini 3.1 Pro · Grok 4.3$2.50–$12.00
PremiumGPT-5.5 · Claude Opus 4.8 · Fable 5$25.00–$50.00
Open-weight models (DeepSeek V4, Nemotron, Kimi K2.6) reach top-10 AA scores at a fraction of closed-flagship output cost

Specialized Performance Highlights

Speed & Latency
Grok 4.3
Fastest AA task completion (~1.5 min) at $1.25 / $2.50 — best wall-clock value in the top tier
GPT-5.5 Instant & Flash-class siblings
GPT-5.5 Instant (May 5) is the new ChatGPT default for low-latency work. Gemini 3.5 Flash and Claude Haiku-class models sit outside the one-per-vendor table — use for interactive and streaming applications
NVIDIA Nemotron 3 Super
Competitive throughput at open-weight cost — strong for high-volume enterprise inference pipelines
Open-Weight Excellence
ModelKey Strength
GLM-5.2AA 51 · SWE-bench Pro 62.1% · 1M ctx · MIT
Gemma 4 12BLaptop-class · encoder-free multimodal · Apache 2.0 · AA 22
Gemma 4 31BAA 29 · multimodal · 256K ctx
DeepSeek V4 ProAA 44 · 1M context · best $/task open weight
MiniMax-M3AA 44 · tied DeepSeek · multimodal agents
Llama 4 Scout10M-token context · corpus-scale tasks
Kimi K2.6 / K2.7-CodeAA 43 · long-horizon coding specialist
GLM-5.2 is the open-weight AA benchmark (51). Gemma 4 12B wins on deployment footprint — not composite score. gpt-oss-120B (AA 24) remains useful for managed API access at $0.30/M.

Model Selection Guide

Peak Intelligence
Claude Opus 4.8Claude Sonnet 5GPT-5.5Gemini 3.5 Flash
Opus 4.8 for peak available AA (56); Sonnet 5 (53) for most production workloads at 40–60% lower cost; GPT-5.5 for Terminal-Bench; Gemini 3.5 Flash when speed matters
Coding & Agents
Claude Sonnet 5Claude Opus 4.8GLM-5.2Gemini 3.5 Flash
Sonnet 5 for daily agentic coding (Terminal-Bench 80.4%); Opus 4.8 for SWE-bench Pro peak; GLM-5.2 for open-weight coding agents; Gemini 3.5 Flash for throughput
Massive Context
Llama 4 ScoutGPT-5.5DeepSeek V4 Pro
Llama 4 Scout (10M tokens, open-weight) for full-codebase tasks; GPT-5.5 (1.05M), DeepSeek V4 Pro (1M), and Gemini 3.1 Pro (1M) for closed and open long-document pipelines
Cost Optimization
DeepSeek V4 FlashNVIDIA NemotronGemini 3.5 Flash
Nemotron and DeepSeek V4 Flash for lowest per-token cost; Gemini 3.5 Flash for frontier-class speed at mid-tier pricing ($9/M output)
Self-Hosting
GLM-5.2Gemma 4 12BLlama 4 ScoutDeepSeek V4 Pro
GLM-5.2 for best open-weight AA (51) with 1M ctx; Gemma 4 12B for laptop/single-GPU multimodal; Llama 4 Scout for 10M context; DeepSeek V4 Pro for best $/task
Agentic Engineering
GLM-5.2Claude Sonnet 5
GLM-5.2 for open-weight long-horizon coding (SWE-bench Pro 62.1%, 1M ctx); Sonnet 5 for closed-model agentic engineering at mid-tier pricing

Industry Impact & Future Trends (2026)

The 2026 LLM landscape is defined by AA Index v4.1's agentic re-weighting, regulatory friction on frontier releases, and a collapsing mid-tier:

Coding & Agents
GLM-5.2 leads open-weight SWE-bench Pro (62.1%). Opus 4.8 leads closed (69.2%). Sonnet 5 is the new agentic default — Terminal-Bench 2.1 80.4%. Kimi K2.7-Code targets repository-scale work.
Regulatory & Access
Fable/Mythos 5 suspended June 12 under export-control review; staged return underway. GPT-5.6 Sol preview limited to vetted partners. Gemini 3.5 Pro delayed to July for quality fixes. Government pre-release review becoming a recurring friction point.
Open-Weight & Local AI
Gemma 4 12B (June 3) brings encoder-free multimodal to 16GB laptops under Apache 2.0. GLM-5.2 adds 1M ctx at AA 51. The open-weight story is now coding-agents-first, not just parameter efficiency.

Conclusion

The July 2026 landscape shifted in four weeks. AA Index v4.1 re-baselined scores with heavier agentic weighting — compare within v4.1, not against June's v4.0 numbers. Anthropic launched Sonnet 5 as the default mid-tier agent, while Fable/Mythos 5 briefly topped the index at AA 60 before a government-mandated suspension. Z.ai's GLM-5.2 became the leading open-weight model (AA 51, SWE-bench Pro 62.1%). Google's Gemma 4 12B redefined laptop-class multimodal deployment. OpenAI previewed GPT-5.6 Sol under restricted access; Gemini 3.5 Pro is delayed to July.

Strategic Takeaway (2026)

Use the right leaderboard for the question. Vendor diversity (Table 1) answers "which companies matter?" Model ranking (Table 2) answers "which specific model should I deploy?" Match workload to model: peak available AA → Opus 4.8; daily agentic coding → Sonnet 5; open-weight coding agents → GLM-5.2; laptop multimodal → Gemma 4 12B; speed + value API → Gemini 3.5 Flash; massive context → Llama 4 Scout (10M) or GLM-5.2 (1M); Mythos-class (if accessible) → Fable/Mythos 5. Regulatory access limits now matter as much as benchmark scores.

Looking ahead: Gemini 3.5 Pro and GPT-5.6 general availability will reshuffle both tables. AA v4.1's agentic focus means Terminal-Bench and GDPval-AA matter more than static knowledge benchmarks. Expect another refresh when those models ship broadly — likely within weeks of this update.