LLM Landscape 2026: Intelligence Leaderboard and Model Guide

An August 2026 snapshot of the frontier LLM landscape, with two Top 25 leaderboards: one model per vendor (company diversity) and a pure model ranking by AA Intelligence Index v4.1 score (multiple models per lab allowed). Since July 1, Anthropic shipped Claude Opus 5 (AA 61, Jul 24) — the new AA leader, priced the same as the Opus 4.8 it replaces. OpenAI's GPT-5.6 Sol reached general availability (Jul 9, ending its June 26 preview) and had its API price cut twice, most recently -20%/-33% on Aug 21. xAI shipped two Grok generations in five weeks — Grok 4.5 (Jul 8) then Grok 4.6 (Aug 12, AA 61, tying Opus 5). Google shipped Gemini 3.7 Flash (Aug 13, AA 56) while Gemini 3.5 Pro remains unreleased three months after its May announcement. Most strikingly, open-weight models nearly closed the gap: Moonshot AI's Kimi K3 (AA 60, Jul 16) and Z.ai's GLM-5.3 (AA 60, Aug 14) both sit just one point behind Opus 5 — the tightest the open/closed gap has ever been on this index. Late August then reset the cost curve: Z.ai unmasked the viral OpenRouter stealth model Ox Alpha as GLM-5.3-Flash (Aug 26, AA 57, $0.15/$0.50), Alibaba shipped Qwen3.8-Flash-Next (AA 56) as a Qwen4 architecture preview plus dense Qwen3.8-27B (AA 52), and DeepSeek's dated V4 refreshes plus an experimental Flash Vision sibling round out a genuinely frontier-moving two months.

Leaderboard Methodology

Both tables use the AA (Artificial Analysis) Intelligence Index v4.1, which aggregates nine agentic-weighted evaluations — GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR — into a single normalized integer. AA now labels that same nine-eval suite v4.1.1; newly scored models (GLM-5.3-Flash, Qwen3.8-Flash-Next, Qwen3.8-27B) are taken from the current listing. Previously published integers in these tables are left unchanged unless AA re-ran a specific checkpoint — compare within v4.1/v4.1.1, not against pre-v4.1 numbers. Scores shown use each model's highest published effort tier (typically max / xhigh). Context windows use comma-separated token counts; missing public data is "—". Pricing is per million tokens (input / output).

Table 1 — Top 25 by vendor: exactly one representative per company — the provider's newest clearly superior flagship or best overall product — to surface geographic and corporate diversity. Table 2 — Top 25 by model: pure AA Index ranking; multiple entries from Anthropic, OpenAI, Google, and others are expected. Models with limited or suspended API access (Fable 5, Mythos 5) are included with availability notes because they materially affect the competitive picture.

Top 25 LLMs by Vendor — Company Diversity Leaderboard (August 2026)

RankModelCapability Index
(AA Index)
Context Window
(tokens)
Input Cost
($/M tokens)
Output Cost
($/M tokens)
Notes
1
Claude Opus 5
Anthropic
611,000,000$5.00$25.00New Jul 24 New AA leader; succeeds Opus 4.8 at the same price; comes close to Fable 5's intelligence at half the cost
2
Grok 4.6
xAI
61500,000$2.00$6.00New Aug 12 Ties Claude Opus 5; two Grok generations shipped in five weeks (4.5 → 4.6); built for long-running agents
3
GLM-5.3
Z.ai
601,000,000$1.40$4.40New Aug 14 Open Weight Ties Kimi K3 for #1 open-weight, 1 pt behind Opus 5; same 753B base as GLM-5.2, gains from post-training alone; weights staged ~Aug 28. Flash sibling (AA 57, $0.15/$0.50) shipped Aug 26
4
Kimi K3
Moonshot AI
601,049,000$3.00$15.00Open Weight Jul 16 launch, weights opened Jul 26; 2.8T params, native multimodal — largest open-weight model to date
5
GPT-5.6 Sol
OpenAI
591,050,000$4.00$20.00GA since Jul 9 (June 26 preview ended); price cut -20%/-33% on Aug 21; most token-efficient frontier model (~15k tokens/task)
6
Qwen3.8 Max
Alibaba
581,000,000New Aug 3 Open Weight 2.4T params (95B active); open weights Aug 12 — first Qwen-Max-class model released open. Efficiency siblings: Flash-Next (AA 56) and dense 27B (AA 52)
7
Gemini 3.7 Flash
Google
561,000,000$0.75$3.75New Aug 13 Fastest reasoning model on the market (~340 t/s); Gemini 3.5 Pro still unreleased, 3+ months after its May announcement — Gemini 4 now in pretraining
8
DeepSeek V4 Pro-0813
DeepSeek AI
531,000,000$2.19$8.76Open Weight Dated GA checkpoint; +9 pts vs the original V4 Pro; V4 Flash-0731 and an experimental Flash Vision sibling cover budget/multimodal
9
MiniMax-M3
MiniMax
441,000,000Open Weight Strong SWE-bench Verified (~80.5%)
10
MiMo-V2.5-Pro
Xiaomi
421,000,000Successor to MiMo-V2-Pro; pricing not publicly disclosed
11
NVIDIA Nemotron 3 Super 120B
NVIDIA
251,000,000$0.30$0.75Open Weight Enterprise self-hosting value tier; Mamba-2/MoE architecture
12
Mistral Large 3
Mistral
16256,000$0.50$1.50Open Weight EU-based; Apache 2.0; v4.1 agentic re-weighting lowered composite vs v4.0
13
Nova Premier
Amazon
131,000,000$2.50$12.50Hyperscaler representative; deep AWS integration
14
Llama 4 Scout
Meta
1010,000,000Open Weight Context-window outlier; 10M tokens for corpus-scale self-hosting
15
Command A
Cohere
8256,000$2.50$10.00Enterprise RAG and tool-use focus
16
Solar Pro 2
Upstage
831B Korean frontier; strong regional/language performance; v4.1 composite lower than legacy v4.0 ranking
17
ERNIE 4.5 300B A47B
Baidu
Best verifiable ERNIE-family public entry; AA v4.1 pending
18
Granite 4.0 H Small
IBM
5Open Weight Enterprise governance and open-deployment focus
19
Jamba 1.7 Large
AI21
5Hybrid SSM/Transformer architecture for long-input efficiency
20
Yi-Lightning
01.AI
Vendor-diversity slot; public AA v4.1 score not yet published
21
Sonar Reasoning Pro
Perplexity
128,000Search-augmented reasoning API; AA v4.1 pending
22
Reka Flash 3
Reka
Multimodal agentic model; AA v4.1 pending
23
Hunyuan-A13B-Instruct
Tencent
Chinese hyperscaler representative; AA v4.1 pending
24
Stable LM 2 12B
Stability AI
Open Weight Community/open-deployment slot; AA v4.1 pending
25
Evo-Ukiyoe
Sakana AI
Evolutionary-model research lab; specialized rather than general-purpose frontier

Top 25 Models by AA Index v4.1 — Pure Capability Leaderboard (August 2026)

This table ranks the twenty-five highest-scoring models on the AA Index v4.1 regardless of vendor — expect multiple Anthropic, OpenAI, and Google entries. Effort tiers are max / xhigh unless noted.

RankModelCapability Index
(AA v4.1)
Context Window
(tokens)
Input Cost
($/M tokens)
Output Cost
($/M tokens)
Notes
1
Claude Opus 5
Anthropic
611,000,000$5.00$25.00New Jul 24 New AA leader; succeeds Opus 4.8 two months after it shipped, at the same price
2
Grok 4.6 (high)
xAI
61500,000$2.00$6.00New Aug 12 Ties Opus 5; ~4x cheaper output than Opus 5; long-running agent and visual-work focus
3
Claude Fable 5
Anthropic
601,000,000$10.00$50.00Limited access Mythos-class with safety classifiers; suspended June 12, staged return expected; Opus 5 fallback on blocked queries
4
GLM-5.3
Z.ai
601,000,000$1.40$4.40New Aug 14 Open Weight Ties Kimi K3 for open-weight #1; same 753B/40B-active base as GLM-5.2; weights staged ~Aug 28. Flash sibling (AA 57) is the cost play, not a replacement
5
Kimi K3
Moonshot AI
601,049,000$3.00$15.00Open Weight Jul 16 launch, weights opened Jul 26; 2.8T params; native multimodal; largest open-weight model to date
6
GPT-5.6 Sol (max)
OpenAI
591,050,000$4.00$20.00GA since Jul 9; most token-efficient frontier model (~15k tokens/task, ~$1.04/task)
7
Qwen3.8 Max
Alibaba
581,000,000New Aug 3 Open Weight 2.4T params (95B active); open weights Aug 12; leads on PaperBench and OSWorld-Verified
8
GLM-5.3-Flash
Z.ai
571,000,000$0.15$0.50New Aug 26 Open Weight Native multimodal; 320B/18B-active; MIT; stealth-tested as Ox Alpha on OpenRouter; ~10× cheaper than GLM-5.3
9
Claude Opus 4.8
Anthropic
561,000,000$5.00$25.00Superseded by Opus 5; SWE-bench Pro 69.2%; still pinned in many production stacks
10
Gemini 3.7 Flash (high)
Google
561,000,000$0.75$3.75New Aug 13 Fastest reasoning model on the market (~340 t/s); intro pricing
11
Qwen3.8-Flash-Next
Alibaba
56262,000New Aug 26 Open Weight Qwen4 architecture preview; 125B/6B-active + 51B n-gram embeddings; 256K native (1M via YaRN)
12
GPT-5.6 Terra (max)
OpenAI
551,050,000Matches GPT-5.5 at roughly half the cost; GA since Jul 9, price cut Jul 30
13
GPT-5.5
OpenAI
551,050,000$5.00$30.00Prior OpenAI flagship; superseded by the GPT-5.6 family
14
Claude Opus 4.7
Anthropic
541,000,000$5.00$25.00Two Opus generations back; still strong for pinned production workflows
15
Grok 4.5
xAI
54500,000$2.00$6.00Jul 8 release, superseded by Grok 4.6 in just 5 weeks; #1 on agentic tool use (τ³-Banking) at launch
16
Claude Sonnet 5
Anthropic
531,000,000$2.00$10.00Default Free/Pro model since Jul 1; Terminal-Bench 2.1 80.4%; intro pricing through Aug 31
17
DeepSeek V4 Pro-0813
DeepSeek AI
531,000,000$2.19$8.76Open Weight Dated GA checkpoint, +9 pts vs the original V4 Pro
18
DeepSeek V4 Flash-0731
DeepSeek AI
521,000,000Open Weight MIT-licensed; 284B total / 13B active MoE; re-post-trained checkpoint, same architecture as the April preview
19
Qwen3.8-27B (xhigh)
Alibaba
52262,000$0.43$2.55New Aug 14 Open Weight Dense 27B; Apache 2.0; native multimodal; laptop-class, ties GPT-5.6 Luna
20
GLM-5.2 (max)
Z.ai
511,000,000$1.40$4.40Open Weight Superseded by GLM-5.3; June 13 GA; SWE-bench Pro 62.1%; MIT license
21
GPT-5.6 Luna (max)
OpenAI
511,050,000Matches/exceeds Gemini 3.5 Flash and GLM-5.2 at lower cost; access expanded to Free/Go users Aug 6
22
Gemini 3.5 Flash (high)
Google
501,000,000$1.50$9.00Superseded by Gemini 3.7 Flash; still GA and widely deployed
23
Claude Sonnet 4.6 (max)
Anthropic
471,000,000$3.00$15.00Superseded by Sonnet 5; still pinned in production agent stacks
24
Gemini 3.1 Pro Preview
Google
461,000,000$2.00$12.00Remains Google's Pro-tier flagship while Gemini 3.5 Pro stays unreleased; multimodal strength
25
MiniMax-M3
MiniMax
441,000,000Open Weight Strong multimodal and agentic scores; SWE-bench Verified ~80.5%

Ox Alpha (also styled 0x Alpha) is not a separate model — Z.ai stealth-tested GLM-5.3-Flash anonymously on OpenRouter and OpenCode from ~Aug 20 before the Aug 26 unmasking. Claude Mythos 5 shares Fable 5's underlying weights and AA score (~60) but is restricted to Project Glasswing partners (cyber/biology safeguards lifted). DeepSeek V4 Flash Vision Exp (Aug 21) is an experimental, API-only vision sibling built on the V4 Flash-0731 backbone — no AA composite score yet, so it stays off the ranked tables; it splits roughly evenly against Claude Opus 4.8 on dedicated visual benchmarks (ALE, ZeroBench) and has no public weights. MiMo-V2.5-Pro, Qwen3.5 397B, NVIDIA Nemotron 3 Super, Grok 4.3/4.20, GPT-5.4, o3/o4-mini, gpt-oss-120B, Kimi K2.6/K2.7-Code, and the Gemma 4 family fell out of this cycle's top 25 amid the new arrivals but remain relevant lower-cost options — see the model selector for full listings.

Key Takeaways

New Leader, and the Open-Weight Gap Nearly Closes
Claude Opus 5 (Jul 24, AA 61) is the new leader among generally-available models, just one point behind Fable 5's suspended 60. But the real headline is how close open weights got: GLM-5.3 and Kimi K3 (both AA 60) trail Opus 5 by a single point — the tightest the open/closed gap has ever been on this index. Grok 4.6 (AA 61) ties Opus 5 outright at a fraction of the price.
A Frantic Six Weeks for Closed Frontier Models
OpenAI took GPT-5.6 Sol to GA (Jul 9) and cut its price twice, most recently -20%/-33% on Aug 21. xAI shipped Grok 4.5 (Jul 8) and superseded it with Grok 4.6 just five weeks later (Aug 12). Anthropic shipped Opus 5 only two months after Opus 4.8 — the fastest Opus-to-Opus cadence yet, leaving only Haiku without a 5-series refresh. Late August then stacked three efficiency releases in a week: GLM-5.3-Flash, Qwen3.8-Flash-Next, and Qwen3.8-27B.
Open-Weight Surge
Kimi K3 (Jul 16, 2.8T params, native multimodal) became the largest open-weight model yet; GLM-5.3 (Aug 14) matched it a month later using the same 753B base as GLM-5.2 — a pure post-training gain. GLM-5.3-Flash (Aug 26, AA 57) is the first native-multimodal GLM-5, MIT-licensed at 320B/18B-active. Qwen3.8 Max (Aug 3, AA 58) is Alibaba's first Max-class model shipped with open weights; Qwen3.8-Flash-Next (AA 56) previews Qwen4, and dense Qwen3.8-27B (AA 52) matches Luna on a laptop-class file. DeepSeek added a dated Pro refresh (V4 Pro-0813, AA 53) and an MIT-licensed Flash refresh (V4 Flash-0731, AA 52).
Coding & Agentic Leadership
Claude Fable 5 still leads closed-model SWE-bench Pro (80.3%); Opus 5 and Opus 4.8 remain the strongest generally-available closed coders. GLM-5.3 leads open-weight agentic gains — its GDPval-AA v2 Elo jumped 246 points to 1,770, second only to Opus 5 (1,855) and ahead of Kimi K3 (1,668). GLM-5.3-Flash matches that GDP work at 1,773 Elo for a tenth of the price. Grok 4.5 held the top spot on agentic tool use (τ³-Banking) at its launch, a distinction Grok 4.6 builds on for long-running agents.
Context-Window Outlier
Llama 4 Scout still pushes open-weight context to 10,000,000 tokens — enabling full-codebase and corpus-scale analysis in a single pass. Most new closed and open-weight flagships this cycle (Opus 5, Sol, GLM-5.3, GLM-5.3-Flash, Kimi K3, Qwen3.8 Max) cluster at ~1,000,000 tokens; Grok 4.6 is a notable exception at 500,000, and Qwen3.8-Flash-Next / Qwen3.8-27B ship at 262K native (1M via YaRN).
Cost-Efficient Frontier
GLM-5.3-Flash ($0.15/$0.50) is the new cost Pareto: AA 57 at roughly a tenth of GLM-5.3 and a fiftieth of Opus 5 output. Gemini 3.7 Flash's intro pricing ($0.75/$3.75) remains the fastest frontier-class API. Grok 4.6 ties the #1 AA score at $2/$6 — roughly a quarter of Opus 5's output cost. GPT-5.6 Sol's Aug 21 cut to $4/$20 and Sonnet 5's intro pricing ($2/$10 through Aug 31) keep closed models competitive on price, but DeepSeek V4 Pro-0813 ($2.19/$8.76) and Nemotron ($0.30/$0.75) remain the open-weight volume anchors below Flash.

Key Performance Metrics

Task-Specific Leaders
ModelBenchmark Leadership
Claude Opus 5New AA leader (61) · available today · same price as Opus 4.8
Grok 4.6Ties Opus 5 (61) · $2/$6 · long-running agents
Claude Fable 5SWE-bench Pro 80.3% · limited access · Mythos-class
GLM-5.3 / Kimi K3Tied open-weight leaders (60) · 1 pt behind Opus 5
GPT-5.6 SolGA Jul 9 · most token-efficient frontier model · AA 59
Qwen3.8 MaxOpen weights Aug 12 · leads PaperBench, OSWorld-Verified
GLM-5.3-FlashAA 57 · $0.15/$0.50 · native multimodal · former Ox Alpha
Context Window Champions
ModelTokens
Llama 4 Scout10,000,000
Opus 5 · GLM-5.3 · GLM-5.3-Flash · Kimi K3 · Qwen3.8 Max1,000,000–1,049,000
GPT-5.6 Sol · DeepSeek V4 Pro-0813 · Nemotron1,000,000–1,050,000
Qwen3.8-Flash-Next · Qwen3.8-27B262,000 (1M YaRN)
Grok 4.6500,000
10M tokens fits entire codebases; 500K–1M+ covers most legal corpora and research archives in one pass
Cost Efficiency
TierModelsOutput $/M
Best ValueGLM-5.3-Flash · Nemotron · Gemini 3.7 Flash~$0.50–$3.75
Mid-RangeGrok 4.6 · Sonnet 5 · Qwen3.8-27B$2.55–$10.00
FlagshipGPT-5.6 Sol · Gemini 3.1 Pro$12.00–$20.00
PremiumClaude Opus 5 · GPT-5.5 · Fable 5$25.00–$50.00
Open-weight models (GLM-5.3-Flash, GLM-5.3, Kimi K3, Qwen3.8-27B) now reach top-20 AA scores at a fraction of closed-flagship output cost

Specialized Performance Highlights

Speed & Latency
Gemini 3.7 Flash
~340 tokens/sec — the fastest reasoning model on the market, at $0.75/$3.75 intro pricing
GLM-5.3-Flash (cost, not speed)
~50 t/s — "Flash" here means price ($0.15/$0.50) and 18B-active MoE, not tokens/sec. AA 57 at a tenth of GLM-5.3
GPT-5.6 Luna & Qwen3.8-Flash-Next
Luna is the low-latency GPT-5.6 tier for Free/Go ChatGPT users. Qwen3.8-Flash-Next (~74 t/s) is the open-weight efficiency preview of Qwen4
Open-Weight Excellence
ModelKey Strength
GLM-5.3 / Kimi K3AA 60 tied #1 open-weight · 1 pt behind Opus 5
GLM-5.3-FlashAA 57 · $0.15/$0.50 · 320B/18B · MIT · native multimodal
Qwen3.8 MaxAA 58 · 2.4T params · first open Qwen-Max model
Qwen3.8-Flash-NextAA 56 · Qwen4 preview · 125B/6B active
Qwen3.8-27BAA 52 · dense 27B · Apache 2.0 · laptop-class
DeepSeek V4 Pro-0813AA 53 · dated GA checkpoint · best $/task open weight
Llama 4 Scout10M-token context · corpus-scale tasks
GLM-5.3 and Kimi K3 are the open-weight AA benchmark (60). GLM-5.3-Flash (57) and Qwen3.8-27B (52) are the new efficiency and laptop-class stories. gpt-oss-120B (AA 24) remains useful for managed API access at $0.30/M.

Model Selection Guide

Peak Intelligence
Claude Opus 5Grok 4.6GPT-5.6 SolClaude Sonnet 5
Opus 5 for peak available AA (61) at Opus 4.8 pricing; Grok 4.6 ties it at a quarter of the cost; GPT-5.6 Sol for the most token-efficient closed flagship; Sonnet 5 for most production workloads at 60–80% lower cost
Coding & Agents
Claude Sonnet 5Claude Opus 5GLM-5.3Kimi K3
Sonnet 5 for daily agentic coding (Terminal-Bench 80.4%); Opus 5 for peak closed-model coding; GLM-5.3 for open-weight agentic work (GDPval-AA v2 Elo 1,770); Kimi K3 for open-weight multimodal agents
Massive Context
Llama 4 ScoutGPT-5.6 SolQwen3.8 Max
Llama 4 Scout (10M tokens, open-weight) for full-codebase tasks; GPT-5.6 Sol (1.05M) and Qwen3.8 Max (1M, open-weight) for closed and open long-document pipelines
Cost Optimization
GLM-5.3-FlashNVIDIA NemotronGemini 3.7 Flash
GLM-5.3-Flash ($0.50/M) for frontier-adjacent intelligence at flash price; Nemotron ($0.75/M) for lowest open-weight per-token cost; Gemini 3.7 Flash ($3.75/M) for the fastest frontier-class speed at budget pricing
Self-Hosting
GLM-5.3GLM-5.3-FlashQwen3.8-27BLlama 4 Scout
GLM-5.3 for best open-weight AA (60); GLM-5.3-Flash (320B/18B, MIT) for cheaper self-host; Qwen3.8-27B for laptop-class AA 52; Llama 4 Scout for 10M context; Gemma 4 12B for single-GPU multimodal
Agentic Engineering
GLM-5.3-FlashGLM-5.3Grok 4.6
GLM-5.3-Flash matches GLM-5.3 on GDPval-AA v2 (Elo 1,773) at flash price; GLM-5.3 for peak open-weight agentic work; Grok 4.6 for closed-model long-running agents; Sonnet 5 for agentic engineering at mid-tier pricing

Industry Impact & Future Trends (2026)

The 2026 LLM landscape is defined by the fastest release cadence yet across both closed and open models, and a closing open/closed intelligence gap:

Coding & Agents
Claude Fable 5 leads closed-model SWE-bench Pro (80.3%). GLM-5.3's GDPval-AA v2 Elo jump (+246 pts to 1,770) is the largest agentic gain of the cycle, trailing only Opus 5; GLM-5.3-Flash matches it at 1,773 Elo. Grok 4.5 topped agentic tool use at launch; Grok 4.6 builds on that for long-running agents.
Regulatory & Access
Fable/Mythos 5 remain suspended since June 12 under export-control review. Gemini 3.5 Pro has missed at least four internal deadlines since its May announcement; Google has begun pretraining Gemini 4 instead. GPT-5.6, by contrast, moved from a government-gated preview to full GA and two price cuts within eight weeks.
Open-Weight & Local AI
Kimi K3 and GLM-5.3 both hit AA 60 — one point off the new closed-model leader. GLM-5.3-Flash (AA 57, MIT, native multimodal) and Qwen3.8-Flash-Next (AA 56, Qwen4 preview) are the late-August efficiency wave; Qwen3.8-27B (AA 52) is laptop-class frontier. DeepSeek keeps iterating on dated checkpoints (V4 Pro-0813, V4 Flash-0731) and previewed an experimental vision sibling. The open-weight story is now genuinely frontier-adjacent, not just a value play.

Conclusion

The August 2026 update covers the busiest stretch this landscape has tracked. Anthropic shipped Claude Opus 5 (AA 61, Jul 24) as the new leader among generally-available models — just two months after Opus 4.8. OpenAI took GPT-5.6 Sol to general availability (Jul 9) and cut its price twice since. xAI shipped and superseded an entire model generation (Grok 4.5 → Grok 4.6) in five weeks, with 4.6 tying Opus 5's AA score outright. Google shipped Gemini 3.7 Flash while Gemini 3.5 Pro remains unreleased three months after its announcement. Late August then unmasked the OpenRouter stealth model Ox Alpha as GLM-5.3-Flash (AA 57, $0.15/$0.50) and added Alibaba's Qwen3.8-Flash-Next (AA 56) and dense Qwen3.8-27B (AA 52). Open weights had their best stretch yet: Kimi K3 and GLM-5.3 both reached AA 60, a single point off the new closed-model leader, while Qwen3.8 Max became Alibaba's first open Max-class model and DeepSeek shipped two more dated V4 checkpoints plus an experimental vision variant.

Strategic Takeaway (2026)

Use the right leaderboard for the question. Vendor diversity (Table 1) answers "which companies matter?" Model ranking (Table 2) answers "which specific model should I deploy?" Match workload to model: peak available AA → Opus 5 or Grok 4.6; daily agentic coding → Sonnet 5; open-weight agentic coding → GLM-5.3 or Kimi K3; open-weight flash cost → GLM-5.3-Flash; laptop-class open weight → Qwen3.8-27B; largest open Qwen model → Qwen3.8 Max; speed + value API → Gemini 3.7 Flash; massive context → Llama 4 Scout (10M); Mythos-class (if accessible) → Fable/Mythos 5. The open/closed intelligence gap is now down to a single point at the top — deployment constraints (licensing, region, self-hosting) matter as much as raw benchmark scores.

Looking ahead: Gemini 3.5 Pro (or a Gemini 4/3.7 Pro successor) and GLM-5.3's pending open-weight release (~Aug 28) are still the two most consequential unresolved threads — GLM-5.3-Flash weights are already public. Watch for whether Kimi K3 or GLM-5.3 pulls ahead once both settle, whether Qwen3.8-Flash-Next's Qwen4 architecture graduates to a production Max-class successor, and whether DeepSeek's Flash Vision Exp graduates from experimental to a full open-weight release. Expect another refresh within a few weeks given this cycle's pace.