What changed on the public AI leaderboards in the 9 September 2026 refresh of the thinQit Index, and what it means for the rankings.
What changed
- Claude Mythos Preview (Anthropic) entered the frontier leaderboard at #1 with a thinQit Index of 91.5.
- Seed-2.0-Mini (ByteDance Seed) entered the local leaderboard at #2 with a thinQit Index of 71.7.
- Muse Spark 1.3 (Meta) entered the frontier leaderboard at #4 with a thinQit Index of 77.6.
- Seed-2.0-Lite (ByteDance Seed) entered the local leaderboard at #5 with a thinQit Index of 69.5.
- Seed 2.0 Pro (ByteDance Seed) entered the frontier leaderboard at #6 with a thinQit Index of 76.5.
- Inkling Small (Thinking Machines) entered the local leaderboard at #8 with a thinQit Index of 67.4.
- Muse Glimmer-30B (Meta) entered the local leaderboard at #9 with a thinQit Index of 65.9.
- GLM-5.2 (Z.ai) entered the frontier leaderboard at #13 with a thinQit Index of 72.0.
Why it matters
It enters with 2 of 7 capabilities covered. Its lead over #2 Claude Fable 5 is 11.6 Index points.
Also refreshed today: Ling 3.0 Flash: AIME 2026 93.2% · Ling 3.0 Flash: Blended price 0.03 · Claude Mythos Preview: SWE-bench Verified 93.9% · Claude Mythos Preview: GPQA Diamond 94.6% · Claude Mythos Preview: Humanity's Last Exam 64.7% · Claude Fable 5: AA Intelligence Index 49.7 · Claude Fable 5: Output speed 68.13 · Seed 2.1 Pro: Humanity's Last Exam 55.7% and 128 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Mythos Preview | Anthropic | 91.5 | — |
| 2 | Claude Fable 5 | Anthropic | 79.9 | −0.2 |
| 3 | Claude Fable 5.1 | Anthropic | 78.1 | −0.3 |
| 4 | Muse Spark 1.3 | Meta | 77.6 | — |
| 5 | Claude Opus 5 | Anthropic | 76.5 | −0.3 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 72.1 |
| 2 | Seed-2.0-Mini | — | 71.7 |
| 3 | GLM-5.3-Flash | 321.3B (32B active) | 69.9 |
| 4 | DeepSeek V4 Flash | 304.2B (24B active) | 69.9 |
| 5 | Seed-2.0-Lite | — | 69.5 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards each morning, recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
