What changed on the public AI leaderboards in the 10 October 2026 refresh of the thinQit Index, and what it means for the rankings.
What changed
- Seed 2.1 Pro lost 12.0 Index points (79.6 → 67.6) after new results: LMArena WebDev new at 1519.
- Hy4 preview moved up from #29 to #1 on the local leaderboard, passing GLM-5.3, Seed-2.0-Mini, DeepSeek V4 Flash, GLM-5.3-Flash, Seed-2.0-Lite, Qwen3.8 2.4T-A95B, Qwen3.8-Flash-Next, Muse Glimmer-30B, Qwen3.8 27B, Qwen3.6 27B, Ling 3.0 Flash, Qwen3.5 397B-A17B, Qwen3.6 35B-A3B, Inkling Small, Qwen3.5 122B-A10B, Gemma 4 31B, gpt-oss-120B, Gemma 4 26B-A4B, gpt-oss-20B, Qwen3.5 9B, Gemma 4 12B, Qwen3.5 4B, Nemotron 3.5 Lightning 30B-A3B, Ministral 3 14B, Granite 4.2 8B, Llama 4 Scout, Gemma 4 E4B and Llama 3.3 70B (Index 79.4 → 78.9).
- GLM-5.3 slipped from #1 to #2 on the local leaderboard, passed by Seed-2.0-Mini (Index 71.9 → 72.2).
- Hy4 preview moved up from #26 to #2 on the frontier leaderboard, passing Claude Fable 5, Claude Fable 5.1, Seed 2.0 Pro, Claude Opus 5, Muse Spark 1.3, Kimi K3, GPT-5.6 Sol, Qwen3.8 Max, GPT-6 Astra, DeepSeek V4 Pro, GPT-5.6 Terra, GLM-5.3, DeepSeek V4.1 Flash, GLM-5.2, Gemini 3.7 Flash, Gemini 3.8 Flash, Muse Spark 1.2, Claude Sonnet 5, Gemini 3.1 Pro, Grok 4.6, Inkling, Kimi K2.7 Code, Mistral Medium 3.5 and Seed 2.1 Pro (Index 79.4 → 78.9).
- Claude Fable 5 slipped from #2 to #4 on the frontier leaderboard, passed by Claude Fable 5.1 and Seed 2.0 Pro (Index 78.4 → 78.0).
- Seed-2.0-Mini slipped from #2 to #3 on the local leaderboard, passed by DeepSeek V4 Flash (Index 71.7 → 71.7).
- DeepSeek V4 Flash slipped from #3 to #4 on the local leaderboard, passed by GLM-5.3-Flash (Index 71.0 → 70.8).
- Inkling Small lost 2.5 Index points (57.8 → 55.3) after new results: LMArena WebDev new at 1409; Output speed 142.34 → 127.31.
Why it matters
LMArena WebDev feeds the Coding capability, which carries 25 % of the Index. The gap to #21 Grok 4.6 is 0.8 Index points. Behind it, #23 Gemini 3.1 Pro is 0.1 points back.
Also refreshed today: Claude Opus 5.5: Output speed 97.28 · Claude Opus 5.5: LMArena WebDev 1813 · Hy4 preview: LMArena WebDev 1632 · Claude Fable 5.1: Output speed 71.89 · Claude Fable 5.1: LMArena WebDev 1744 · Claude Fable 5: LMArena WebDev 1625 · Muse Spark 1.3: Output speed 224.53 · Muse Spark 1.3: LMArena WebDev 1657 and 60 more.
The leaderboard today
| # | Model | Lab | Index | Δ day |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 | Anthropic | 80.8 | +2.4 |
| 2 | Hy4 preview | Tencent | 78.9 | −0.5 |
| 3 | Claude Fable 5.1 | Anthropic | 78.5 | +1.6 |
| 4 | Claude Fable 5 | Anthropic | 78.0 | −0.4 |
| 5 | Seed 2.0 Pro | ByteDance Seed | 76.5 | 0.0 |
| # | Model | Params | Index |
|---|---|---|---|
| 1 | Hy4 preview | 780B (40B active) | 78.9 |
| 2 | GLM-5.3 | 753.3B (40B active) | 72.2 |
| 3 | Seed-2.0-Mini | — | 71.7 |
| 4 | DeepSeek V4 Flash | 304.2B (24B active) | 70.8 |
| 5 | GLM-5.3-Flash | 321.3B (32B active) | 70.4 |
How we measure
The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Every day. A GitHub Actions job re-reads the public leaderboards at midnight (Amsterdam time), recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.
Why does a model show a provisional score?
A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
