What changed on the public AI leaderboards in the 3 September 2026 refresh of the thinQit Index, and what it means for the rankings.
What changed
- GLM-5.3-Flash slipped from #2 to #3 on the local leaderboard, passed by Qwen3.8-Flash-Next (Index 73.6 → 71.0).
- DeepSeek V4 Pro slipped from #8 to #10 on the frontier leaderboard, passed by Qwen3.8 Max and Grok 4.6 (Index 75.1 → 73.0).
- Qwen3.8 Max slipped from #9 to #11 on the frontier leaderboard, passed by Grok 4.6 and GPT-5.6 Terra (Index 74.8 → 72.7).
- Artificial Analysis is being read again after an outage; its benchmarks refreshed today, so scores it feeds moved for that reason rather than because a leaderboard changed.
- LiveBench is being read again after an outage; its benchmarks refreshed today, so scores it feeds moved for that reason rather than because a leaderboard changed.
Why it matters
GLM-5.3-Flash now leads Qwen3.8-Flash-Next by 3.7 points on Coding, the widest gap between the two. Qwen3.8-Flash-Next still wins Agentic and Speed & price. The gap to #2 DeepSeek V4 Flash is 0.4 Index points. Behind it, #4 Qwen3.8 2.4T-A95B is 0.5 points back.
Also refreshed today: Claude Fable 5.1: Output speed 69.33 · Claude Fable 5: Output speed 59.44 · Claude Fable 5: LiveBench Reasoning 91.69% · Claude Fable 5: LiveBench Coding 86.38% · Claude Fable 5: LiveBench Agentic Coding 66.06% · Claude Opus 5: Output speed 48.05 · Claude Opus 5: LiveBench Reasoning 91.21% · Claude Opus 5: LiveBench Coding 81.44% and 79 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 | Anthropic | 81.0 | 0.0 |
| 2 | Claude Fable 5 | Anthropic | 80.8 | 0.0 |
| 3 | Claude Opus 5 | Anthropic | 77.6 | 0.0 |
| 4 | Gemini 3.7 Flash | 75.6 | 0.0 | |
| 5 | Kimi K3 | Moonshot AI | 75.2 | 0.0 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 73.4 |
| 2 | DeepSeek V4 Flash | 304.2B (24B active) | 71.4 |
| 3 | GLM-5.3-Flash | 321.3B (32B active) | 71.0 |
| 4 | Qwen3.8 2.4T-A95B | 2400B (95B active) | 70.5 |
| 5 | Qwen3.8-Flash-Next | 177B (12B active) | 69.2 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards each morning, recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
