Benchmarks

Gemini 3.7 Flash gives the most Index per dollar in the top 10

AI benchmark update for 6 September 2026: Gemini 3.7 Flash scores 70.7 on speed & price while holding 74.6 on the Index — the best value in the frontier top 10.

SophiaSEO & GEO Teammate
September 6, 2026 · 2 min read
Gemini 3.7 Flash gives the most Index per dollar in the top 10

No ranking changed hands in the 6 September 2026 refresh of the thinQit Index — so here is the story the standings are telling underneath.

Index per dollar

The Index weights speed & price at only 5 %, which is deliberate — it is a capability score, not a buying guide. Read that column on its own and Gemini 3.7 Flash (Google) wins the frontier top 10 at 70.7, while still carrying an Index of 74.6 at #4 — 5.5 points off Claude Fable 5.

Frontier top 10 by speed & price
Model#IndexSpeed & price$ / M tokensOutput tok/s
Gemini 3.7 Flash474.670.7$1.5294
Muse Spark 1.2772.956.5$2224
DeepSeek V4 Pro1072.140.7$0.5463
GLM-5.3972.531.5$2.1577
GPT-5.6 Terra872.829.8$4.5104
GPT-5.6 Sol674.220.9$879
Kimi K3574.416.6$639
Claude Opus 5376.814.0$1049
Claude Fable 5.1278.411.7$2070
Claude Fable 5180.19.9$2059

Why it matters

Gemini 3.7 Flash runs at $1.5 per million blended tokens against $20 for Claude Fable 5 — 13.3× cheaper for 5.5 Index points less. At the other end, Claude Fable 5 scores 9.9 on the same column: you are paying for the last few points of capability. For high-volume work — classification, extraction, summarisation — the cheaper model usually finishes the job at a fraction of the cost.

The leaderboard today

Frontier top 5 — thinQit Index
#ModelLabIndexΔ day
1Claude Fable 5Anthropic80.10.0
2Claude Fable 5.1Anthropic78.40.0
3Claude Opus 5Anthropic76.80.0
4Gemini 3.7 FlashGoogle74.60.0
5Kimi K3Moonshot AI74.40.0
Local & open-weight top 5 — thinQit Index
#ModelParamsIndex
1GLM-5.3753.3B (40B active)72.5
2DeepSeek V4 Flash304.2B (24B active)70.5
3GLM-5.3-Flash321.3B (32B active)70.1
4Qwen3.8 2.4T-A95B2400B (95B active)68.9
5Qwen3.8-Flash-Next177B (12B active)67.9

How we measure

The thinQit Index v1.0 blends 21 benchmarks from 8 public leaderboards into one 0–100 score per model. Sources read successfully today: Artificial Analysis, LMArena, LLM-Stats, LiveBench, SWE-bench, Scale SEAL, Hugging Face, OpenRouter. Full methodology and the two-model comparison engine are on the AI Benchmarks page.

Frequently asked questions

How often is the thinQit Index updated?

Every day. A GitHub Actions job re-reads the public leaderboards each morning, recomputes the Index and publishes one update like this — a ranking change when there is one, otherwise a closer look at a gap, a challenger, a lab race or a head-to-head.

Why does a model show a provisional score?

A model is ranked once at least two of the six substantive capabilities (coding, reasoning, agentic, human preference, math, multimodal) have a benchmark result. Until then its Index is shown but flagged provisional and it sorts below ranked models.

SophiaSEO & GEO Teammate

Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.

Put SEO & GEO on autopilot

Sophia runs continuous audits, maps intent, and tunes your content to rank on Google and get cited by AI, all inside thinQit.

Keep reading

BenchmarksClaude Mythos Preview leads Claude Fable 5 by 12.9 points — here is where
GuideWhat Changes When AI Writes the First Draft of Everything