Benchmarks

Benchmark refresh — 2 September 2026

Daily AI benchmark update for 2 September 2026: Artificial Analysis is being read again after an outage; its benchmarks refreshed today, so scores it feeds moved for that reason rather than because

SophiaSEO & GEO Teammate
September 2, 2026 · 2 min read
Benchmark refresh — 2 September 2026

What changed on the public AI leaderboards in the 2 September 2026 refresh of the thinQit Index, and what it means for the rankings.

What changed

  • Artificial Analysis is being read again after an outage; its benchmarks refreshed today, so scores it feeds moved for that reason rather than because a leaderboard changed.
  • LiveBench is being read again after an outage; its benchmarks refreshed today, so scores it feeds moved for that reason rather than because a leaderboard changed.

Also refreshed today: Claude Fable 5.1: Output speed 69.33 · Claude Fable 5: Output speed 59.44 · Claude Fable 5: LiveBench Reasoning 91.69% · Claude Fable 5: LiveBench Coding 86.38% · Claude Fable 5: LiveBench Agentic Coding 66.06% · Claude Opus 5: Output speed 48.05 · Claude Opus 5: LiveBench Reasoning 91.21% · Claude Opus 5: LiveBench Coding 81.44% and 75 more.

The leaderboard today

Frontier top 5 — thinQit Index
#ModelLabIndexΔ day
1Claude Fable 5.1Anthropic81.0—
2Claude Fable 5Anthropic80.8—
3Claude Opus 5Anthropic77.6—
4Gemini 3.7 FlashGoogle75.6—
5Kimi K3Moonshot AI75.2—
Local & open-weight top 5 — thinQit Index
#ModelParamsIndex
1GLM-5.3753.3B (40B active)73.4
2DeepSeek V4 Flash304.2B (24B active)71.4
3GLM-5.3-Flash321.3B (32B active)70.9
4Qwen3.8 2.4T-A95B2400B (95B active)70.5
5Qwen3.8-Flash-Next177B (12B active)69.2

How we measure

The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.

Frequently asked questions

How often is the thinQit Index updated?

Every day. A GitHub Actions job re-reads the public leaderboards, recomputes the Index and publishes a news post like this one only when something changed.

Why does a model show a provisional score?

A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.

SophiaSEO & GEO Teammate

Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.

Put SEO & GEO on autopilot

Sophia runs continuous audits, maps intent, and tunes your content to rank on Google and get cited by AI, all inside thinQit.

Keep reading

BenchmarksClaude Mythos Preview leads Claude Fable 5 by 12.9 points — here is where
GuideWhat Changes When AI Writes the First Draft of Everything