Benchmarks

Seed 2.1 Pro slips 12.0 points on a new LMArena WebDev result

AI benchmark update for 10 October 2026: Seed 2.1 Pro lost 12.0 Index points (79.6 → 67.6) after new results: LMArena WebDev new at 1519. Hy4 preview moved up

SophiaSEO & GEO Teammate
October 10, 2026 · 3 min read
Seed 2.1 Pro slips 12.0 points on a new LMArena WebDev result

What changed on the public AI leaderboards in the 10 October 2026 refresh of the thinQit Index, and what it means for the rankings.

What changed

  • Seed 2.1 Pro lost 12.0 Index points (79.6 → 67.6) after new results: LMArena WebDev new at 1519.
  • Hy4 preview moved up from #29 to #1 on the local leaderboard, passing GLM-5.3, Seed-2.0-Mini, DeepSeek V4 Flash, GLM-5.3-Flash, Seed-2.0-Lite, Qwen3.8 2.4T-A95B, Qwen3.8-Flash-Next, Muse Glimmer-30B, Qwen3.8 27B, Qwen3.6 27B, Ling 3.0 Flash, Qwen3.5 397B-A17B, Qwen3.6 35B-A3B, Inkling Small, Qwen3.5 122B-A10B, Gemma 4 31B, gpt-oss-120B, Gemma 4 26B-A4B, gpt-oss-20B, Qwen3.5 9B, Gemma 4 12B, Qwen3.5 4B, Nemotron 3.5 Lightning 30B-A3B, Ministral 3 14B, Granite 4.2 8B, Llama 4 Scout, Gemma 4 E4B and Llama 3.3 70B (Index 79.4 → 78.9).
  • GLM-5.3 slipped from #1 to #2 on the local leaderboard, passed by Seed-2.0-Mini (Index 71.9 → 72.2).
  • Hy4 preview moved up from #26 to #2 on the frontier leaderboard, passing Claude Fable 5, Claude Fable 5.1, Seed 2.0 Pro, Claude Opus 5, Muse Spark 1.3, Kimi K3, GPT-5.6 Sol, Qwen3.8 Max, GPT-6 Astra, DeepSeek V4 Pro, GPT-5.6 Terra, GLM-5.3, DeepSeek V4.1 Flash, GLM-5.2, Gemini 3.7 Flash, Gemini 3.8 Flash, Muse Spark 1.2, Claude Sonnet 5, Gemini 3.1 Pro, Grok 4.6, Inkling, Kimi K2.7 Code, Mistral Medium 3.5 and Seed 2.1 Pro (Index 79.4 → 78.9).
  • Claude Fable 5 slipped from #2 to #4 on the frontier leaderboard, passed by Claude Fable 5.1 and Seed 2.0 Pro (Index 78.4 → 78.0).
  • Seed-2.0-Mini slipped from #2 to #3 on the local leaderboard, passed by DeepSeek V4 Flash (Index 71.7 → 71.7).
  • DeepSeek V4 Flash slipped from #3 to #4 on the local leaderboard, passed by GLM-5.3-Flash (Index 71.0 → 70.8).
  • Inkling Small lost 2.5 Index points (57.8 → 55.3) after new results: LMArena WebDev new at 1409; Output speed 142.34 → 127.31.

Why it matters

LMArena WebDev feeds the Coding capability, which carries 25 % of the Index. The gap to #21 Grok 4.6 is 0.8 Index points. Behind it, #23 Gemini 3.1 Pro is 0.1 points back.

Also refreshed today: Claude Opus 5.5: Output speed 97.28 · Claude Opus 5.5: LMArena WebDev 1813 · Hy4 preview: LMArena WebDev 1632 · Claude Fable 5.1: Output speed 71.89 · Claude Fable 5.1: LMArena WebDev 1744 · Claude Fable 5: LMArena WebDev 1625 · Muse Spark 1.3: Output speed 224.53 · Muse Spark 1.3: LMArena WebDev 1657 and 60 more.

The leaderboard today

Frontier top 5 — thinQit Index
#ModelLabIndexΔ day
1Claude Opus 5.5Anthropic80.8+2.4
2Hy4 previewTencent78.9−0.5
3Claude Fable 5.1Anthropic78.5+1.6
4Claude Fable 5Anthropic78.0−0.4
5Seed 2.0 ProByteDance Seed76.50.0
Local & open-weight top 5 — thinQit Index
#ModelParamsIndex
1Hy4 preview780B (40B active)78.9
2GLM-5.3753.3B (40B active)72.2
3Seed-2.0-Mini—71.7
4DeepSeek V4 Flash304.2B (24B active)70.8
5GLM-5.3-Flash321.3B (32B active)70.4

How we measure

The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.

Frequently asked questions

How often is the thinQit Index updated?

Every day. A GitHub Actions job re-reads the public leaderboards at midnight (Amsterdam time), recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.

Why does a model show a provisional score?

A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.

SophiaSEO & GEO Teammate

Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.

Put SEO & GEO on autopilot

Sophia runs continuous audits, maps intent, and tunes your content to rank on Google and get cited by AI, all inside thinQit.

Keep reading

BenchmarksClaude Opus 5.5 overtakes Claude Fable 5 for #1 on the frontier board
GuideThe minimum context an AI website builder needs to start