No ranking changed hands in the 8 September 2026 refresh of the thinQit Index — so here is the story the standings are telling underneath.
Frontier watch
Nothing changed hands on the frontier board today, so the question is who is moving underneath it. Claude Fable 5 (Anthropic) holds #1 at 80.1 with Claude Fable 5.1 1.7 points behind. The best-placed challenger from another lab is Muse Spark 1.2 (Meta) at #7, 5.1 points off #2 and 6.8 off the top.
| # | Model | Lab | Index | Δ since 2 Sept | Gap to Muse Spark 1.2 |
|---|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 80.1 | −0.7 | 6.8 |
| 2 | Claude Fable 5.1 | Anthropic | 78.4 | −2.6 | 5.1 |
| 7 | Muse Spark 1.2 | Meta | 73.3 | +0.1 | — |
Why it matters
At the rate of the last 6 days — Muse Spark 1.2 +0.1, Claude Fable 5.1 −2.6 — the gap closes in about 12 days.
Where the gap actually sits: Claude Fable 5.1 leads Muse Spark 1.2 by 18.9 points on Agentic, while Muse Spark 1.2 is already ahead on Speed & price.
Also refreshed today: Claude Fable 5: Output speed 62.17 · Claude Fable 5.1: Output speed 67.63 · Claude Opus 5: Output speed 53.23 · Gemini 3.7 Flash: Output speed 324.77 · Kimi K3: Output speed 43.33 · GPT-5.6 Sol: Output speed 79.88 · Muse Spark 1.2: Output speed 261.99 · GPT-5.6 Terra: Output speed 121.22 and 30 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 80.1 | 0.0 |
| 2 | Claude Fable 5.1 | Anthropic | 78.4 | 0.0 |
| 3 | Claude Opus 5 | Anthropic | 76.8 | 0.0 |
| 4 | Gemini 3.7 Flash | 74.6 | 0.0 | |
| 5 | Kimi K3 | Moonshot AI | 74.4 | 0.0 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 72.5 |
| 2 | DeepSeek V4 Flash | 304.2B (24B active) | 70.4 |
| 3 | GLM-5.3-Flash | 321.3B (32B active) | 70.2 |
| 4 | Qwen3.8 2.4T-A95B | 2400B (95B active) | 68.9 |
| 5 | Qwen3.8-Flash-Next | 177B (12B active) | 68.0 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards each morning, recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
