No ranking changed hands in the 26 September 2026 refresh of the thinQit Index — so here is the story the standings are telling underneath.
Mover of the week
The rankings did not change today, but the Index did. GPT-6 Astra (OpenAI) is up +1.1 points over the past 7 days, from 71.6 to 72.7, and now sits at #10. At the other end, Gemini 3.7 Flash is −3.2 over the same window, from 74.1 to 70.9 at #17.
| # | Model | Lab | Then (19 Sept) | Today | Δ |
|---|---|---|---|---|---|
| 10 | GPT-6 Astra | OpenAI | 71.6 | 72.7 | +1.1 |
| 22 | Inkling | Thinking Machines | 60.0 | 60.9 | +0.9 |
| 7 | GPT-5.6 Sol | OpenAI | 73.9 | 74.4 | +0.5 |
| 11 | GLM-5.3 | Z.ai | 72.0 | 72.4 | +0.4 |
| 3 | Claude Opus 5.5 | Anthropic | 76.6 | 76.8 | +0.2 |
| 17 | Gemini 3.7 Flash | 74.1 | 70.9 | −3.2 |
Why it matters
GPT-6 Astra moved on fresh results: LMArena Vision 1284, LMArena Text 1480, SEAL · HLE 54.8%, Output speed 54.12. Gemini 3.7 Flash re-read on LiveBench Agentic Coding 58.28%, LiveBench Coding 78.89%, LiveBench Reasoning 87.8%, Terminal-Bench 85.77%. A single benchmark rarely moves the Index by more than a point on its own — the weighting spreads it across seven capabilities.
Also refreshed today: Claude Fable 5.1: Output speed 71.5 · Claude Opus 5.5: Output speed 102 · Claude Opus 5: Output speed 68.44 · Muse Spark 1.3: Output speed 241.33 · GPT-5.6 Sol: Output speed 119.43 · Qwen3.8 Max: Output speed 37.28 · GPT-6 Astra: Output speed 54.12 · GLM-5.3: Output speed 91.5 and 32 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 78.6 | 0.0 |
| 2 | Claude Fable 5.1 | Anthropic | 76.8 | 0.0 |
| 3 | Claude Opus 5.5 | Anthropic | 76.8 | +0.2 |
| 4 | Claude Opus 5 | Anthropic | 76.6 | +0.1 |
| 5 | Seed 2.0 Pro | ByteDance Seed | 76.5 | 0.0 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 72.4 |
| 2 | Seed-2.0-Mini | — | 71.7 |
| 3 | DeepSeek V4 Flash | 304.2B (24B active) | 70.9 |
| 4 | GLM-5.3-Flash | 321.3B (32B active) | 69.9 |
| 5 | Seed-2.0-Lite | — | 69.5 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards at midnight (Amsterdam time), recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
