What changed on the public AI leaderboards in the 5 September 2026 refresh of the thinQit Index, and what it means for the rankings.
What changed
- Qwen3.8 2.4T-A95B lost 1.6 Index points (70.5 → 68.9) after new results: AA Intelligence Index 57.7 → 46.7; Output speed 40.22 → 39.79.
- Qwen3.8-Flash-Next lost 1.2 Index points (69.2 → 68.0) after new results: AA Intelligence Index 55.8 → 45.6; Output speed 89.11 → 80.37.
- Claude Sonnet 5 lost 1.1 Index points (71.4 → 70.3) after new results: AA Intelligence Index 55.3 → 45.1; Output speed 78.21 → 80.8.
- Mistral Medium 3.5 lost 1.1 Index points (51.1 → 50.0) after new results: AA Intelligence Index 30.4 → 22.7; Output speed 124.89 → 125.24.
- Qwen3.5 9B lost 1.1 Index points (36.0 → 34.9) after new results: AA Intelligence Index 21.8 → 14.8; Output speed 94.89 → 92.19.
- Claude Fable 5.1 lost 1.0 Index points (79.4 → 78.4) after new results: AA Intelligence Index 65.7 → 56.8; Output speed 70.31 → 71.38.
- Gemini 3.7 Flash lost 1.0 Index points (75.6 → 74.6) after new results: AA Intelligence Index 56 → 45.2; Output speed 297.32 → 302.84.
- DeepSeek V4 Pro lost 1.0 Index points (73.1 → 72.1) after new results: AA Intelligence Index 53.2 → 42.1.
Also refreshed today: Claude Fable 5: AA Intelligence Index 53.2 · Claude Fable 5: Output speed 65.87 · Claude Fable 5.1: AA Intelligence Index 56.8 · Claude Fable 5.1: Output speed 71.38 · Claude Opus 5: AA Intelligence Index 54.1 · Claude Opus 5: Output speed 52.72 · Gemini 3.7 Flash: AA Intelligence Index 45.2 · Gemini 3.7 Flash: Output speed 302.84 and 64 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 80.1 | -0.8 |
| 2 | Claude Fable 5.1 | Anthropic | 78.4 | -1.0 |
| 3 | Claude Opus 5 | Anthropic | 76.8 | -0.8 |
| 4 | Gemini 3.7 Flash | 74.6 | -1.0 | |
| 5 | Kimi K3 | Moonshot AI | 74.4 | -0.8 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 72.5 |
| 2 | DeepSeek V4 Flash | 304.2B (24B active) | 70.6 |
| 3 | GLM-5.3-Flash | 321.3B (32B active) | 70.1 |
| 4 | Qwen3.8 2.4T-A95B | 2400B (95B active) | 68.9 |
| 5 | Qwen3.8-Flash-Next | 177B (12B active) | 68.0 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards, recomputes the Index and publishes a news post like this one only when something changed.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
