No ranking changed hands in the 6 September 2026 refresh of the thinQit Index — so here is the story the standings are telling underneath.
Index per dollar
The Index weights speed & price at only 5 %, which is deliberate — it is a capability score, not a buying guide. Read that column on its own and Gemini 3.7 Flash (Google) wins the frontier top 10 at 70.7, while still carrying an Index of 74.6 at #4 — 5.5 points off Claude Fable 5.
| Model | # | Index | Speed & price | $ / M tokens | Output tok/s |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | 4 | 74.6 | 70.7 | $1.5 | 294 |
| Muse Spark 1.2 | 7 | 72.9 | 56.5 | $2 | 224 |
| DeepSeek V4 Pro | 10 | 72.1 | 40.7 | $0.54 | 63 |
| GLM-5.3 | 9 | 72.5 | 31.5 | $2.15 | 77 |
| GPT-5.6 Terra | 8 | 72.8 | 29.8 | $4.5 | 104 |
| GPT-5.6 Sol | 6 | 74.2 | 20.9 | $8 | 79 |
| Kimi K3 | 5 | 74.4 | 16.6 | $6 | 39 |
| Claude Opus 5 | 3 | 76.8 | 14.0 | $10 | 49 |
| Claude Fable 5.1 | 2 | 78.4 | 11.7 | $20 | 70 |
| Claude Fable 5 | 1 | 80.1 | 9.9 | $20 | 59 |
Why it matters
Gemini 3.7 Flash runs at $1.5 per million blended tokens against $20 for Claude Fable 5 — 13.3× cheaper for 5.5 Index points less. At the other end, Claude Fable 5 scores 9.9 on the same column: you are paying for the last few points of capability. For high-volume work — classification, extraction, summarisation — the cheaper model usually finishes the job at a fraction of the cost.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 80.1 | 0.0 |
| 2 | Claude Fable 5.1 | Anthropic | 78.4 | 0.0 |
| 3 | Claude Opus 5 | Anthropic | 76.8 | 0.0 |
| 4 | Gemini 3.7 Flash | 74.6 | 0.0 | |
| 5 | Kimi K3 | Moonshot AI | 74.4 | 0.0 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | GLM-5.3 | 753.3B (40B active) | 72.5 |
| 2 | DeepSeek V4 Flash | 304.2B (24B active) | 70.5 |
| 3 | GLM-5.3-Flash | 321.3B (32B active) | 70.1 |
| 4 | Qwen3.8 2.4T-A95B | 2400B (95B active) | 68.9 |
| 5 | Qwen3.8-Flash-Next | 177B (12B active) | 67.9 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards each morning, recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
