What changed on the public AI leaderboards in the 12 September 2026 refresh of the thinQit Index, and what it means for the rankings.
What changed
- Muse Spark 1.3 moved up from #7 to #6 on the frontier leaderboard, passing GPT-6 Astra (Index 75.0 → 75.0).
- GPT-6 Astra slipped from #6 to #16 on the frontier leaderboard, passed by Muse Spark 1.3, Gemini 3.7 Flash, GPT-5.6 Sol, Kimi K3, GLM-5.2, Muse Spark 1.2, Gemini 3.8 Flash, GPT-5.6 Terra, GLM-5.3 and DeepSeek V4.1 Flash (Index 75.5 → 71.7).
- Gemini 3.7 Flash moved up from #8 to #7 on the frontier leaderboard, passing Muse Spark 1.3 (Index 74.1 → 74.1).
- GPT-5.6 Sol moved up from #9 to #8 on the frontier leaderboard, passing Gemini 3.7 Flash (Index 73.8 → 73.7).
- Kimi K3 moved up from #10 to #9 on the frontier leaderboard, passing GPT-5.6 Sol (Index 73.8 → 73.5).
- GLM-5.2 moved up from #11 to #10 on the frontier leaderboard, passing Kimi K3 (Index 73.1 → 72.8).
Why it matters
Muse Spark 1.3 now leads GPT-6 Astra by 58.5 points on Speed & price, the widest gap between the two. GPT-6 Astra still wins Reasoning. The gap to #5 Seed 2.0 Pro is 1.5 Index points. Behind it, #7 Gemini 3.7 Flash is 0.9 points back.
Also refreshed today: Claude Fable 5: Output speed 69.98 · Claude Fable 5.1: Output speed 68.16 · Claude Opus 5: Output speed 59.57 · Muse Spark 1.3: Output speed 418.42 · Gemini 3.7 Flash: LMArena Text 1490 · Gemini 3.7 Flash: Output speed 343.31 · GPT-5.6 Sol: Output speed 69.2 · GPT-5.6 Sol: LMArena Text 1482 and 50 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Mythos Preview | Anthropic | 91.5 | 0.0 |
| 2 | Claude Fable 5 | Anthropic | 79.9 | 0.0 |
| 3 | Claude Fable 5.1 | Anthropic | 78.0 | 0.0 |
| 4 | Claude Opus 5 | Anthropic | 76.6 | 0.0 |
| 5 | Seed 2.0 Pro | ByteDance Seed | 76.5 | 0.0 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 72.3 |
| 2 | Seed-2.0-Mini | — | 71.7 |
| 3 | DeepSeek V4 Flash | 304.2B (24B active) | 71.2 |
| 4 | GLM-5.3-Flash | 321.3B (32B active) | 70.4 |
| 5 | Seed-2.0-Lite | — | 69.5 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards each morning, recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
