No ranking changed hands in the 24 September 2026 refresh of the thinQit Index — so here is the story the standings are telling underneath.
Frontier watch
Nothing changed hands on the frontier board today, so the question is who is moving underneath it. Claude Mythos Preview (Anthropic) holds #1 at 91.5 with Claude Fable 5 12.9 points behind. The best-placed challenger from another lab is GPT-6 Astra (OpenAI) at #11, 5.9 points off #2 and 18.8 off the top.
| # | Model | Lab | Index | Δ 7d | Gap to GPT-6 Astra |
|---|---|---|---|---|---|
| 1 | Claude Mythos Preview | Anthropic | 91.5 | 0.0 | 18.8 |
| 2 | Claude Fable 5 | Anthropic | 78.6 | −1.3 | 5.9 |
| 11 | GPT-6 Astra | OpenAI | 72.7 | +1.0 | — |
Why it matters
At the rate of the last 8 days — GPT-6 Astra +1.0, Claude Fable 5 −1.3 — the gap closes in about 21 days.
Where the gap actually sits: Claude Fable 5 leads GPT-6 Astra by 10.4 points on Multimodal, while GPT-6 Astra is already ahead on Speed & price.
Also refreshed today: Claude Fable 5.1: Output speed 63.14 · Claude Opus 5.5: Output speed 83.66 · Claude Opus 5: Output speed 61.33 · Muse Spark 1.3: Output speed 246.99 · GPT-5.6 Sol: Output speed 64.9 · Kimi K3: Output speed 36.67 · Qwen3.8 Max: Output speed 40.2 · GPT-6 Astra: Output speed 55.42 and 34 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Mythos Preview | Anthropic | 91.5 | 0.0 |
| 2 | Claude Fable 5 | Anthropic | 78.6 | 0.0 |
| 3 | Claude Fable 5.1 | Anthropic | 76.7 | −0.1 |
| 4 | Claude Opus 5.5 | Anthropic | 76.6 | 0.0 |
| 5 | Claude Opus 5 | Anthropic | 76.5 | −0.1 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 72.1 |
| 2 | Seed-2.0-Mini | — | 71.7 |
| 3 | DeepSeek V4 Flash | 304.2B (24B active) | 71.0 |
| 4 | GLM-5.3-Flash | 321.3B (32B active) | 69.9 |
| 5 | Seed-2.0-Lite | — | 69.5 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards at midnight (Amsterdam time), recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
