thinQit|AI Benchmarks

Compare language models with the thinQit Index, built from 8 public sources. Explore the wider landscape of image, video, voice, agents and open-weight AI through dated, sourced updates.

Sources8/8 read today
Last refresh25 Sept 2026 · daily
MethodIndex v1.0 weights
Claude Mythos PreviewAnthropic
Best codingBest reasoning
91.5#1 · Index
Claude Fable 5Anthropic
Best multimodalBest human pref.
78.6#2 · Index
Claude Fable 5.1Anthropic
Top 3 overall
76.8#3 · Index
By capability

Where each model wins

Top five per capability. Every chart is drawn to the same 0–100 scale, so a short bar in Math is a short bar.

Coding

AA Coding Index · SWE-bench Verified · LiveCodeBench · LMArena WebDev · LiveBench Coding
02550751009185848281
Claude Mythos PreviewClaude Opus 5.5Claude Fable 5Claude Fable 5.1Claude Opus 5

Reasoning

AA Intelligence Index · GPQA Diamond · Humanity's Last Exam · LiveBench Reasoning · SEAL · HLE
02550751009287848482
Claude Mythos PreviewHy4 previewClaude Opus 5.5Claude Opus 5Claude Fable 5

Math

AA Math Index · AIME 2026
025507510099969592
GLM-5.2InklingSakana NamazuSeed 2.0 Pro

Multimodal

MMMU · LMArena Vision
02550751008481787776
Claude Fable 5Qwen3.8 MaxMuse Spark 1.3Muse Spark 1.2Claude Fable 5.1

Agentic

Terminal-Bench · τ²-bench · LiveBench Agentic Coding · SEAL · MultiChallenge
02550751008581787777
GPT-5.6 TerraClaude Fable 5Claude Fable 5.1DeepSeek V4.1 FlashClaude Opus 5

Human pref.

LMArena Text
02550751008583838181
Claude Fable 5Muse Spark 1.2Claude Fable 5.1Claude Opus 5Muse Spark 1.3
thinQit Index

Frontier leaderboard

One 0–100 score blending coding, reasoning, math, multimodal, agentic and human-preference results. Cells show per-capability scores; the sparkline is Index history.

Refreshed 25 Sept 2026 · 8/8 sources read · refreshes daily
#ModelthinQit IndexHistoryCodingReason.MathMulti.AgenticPref.Score
1Claude Mythos PreviewAnthropic · 1M context————91.5
2Claude Fable 5Anthropic · 1M context—78.6
3Claude Fable 5.1Anthropic · 1M context—76.8▲0.1
4Claude Opus 5.5Anthropic · 1M context———76.6NEW
5Claude Opus 5Anthropic · 1M context—76.5
6Seed 2.0 ProByteDance Seed · 256K context———76.5
7Muse Spark 1.3Meta · 1.049M context—75.6▼0.2
8GPT-5.6 SolOpenAI · 1.05M context—74.0▲0.1
9Kimi K3Moonshot AI · 1.049M context——73.5
10Qwen3.8 MaxAlibaba · Qwen · 1M context—73.0
11GPT-6 AstraOpenAI · 1.05M context—72.7
12GLM-5.3Z.ai · 1.049M context——72.4▲0.3
13Gemini 3.8 FlashGoogle · 1.049M context——72.3▼0.2
14GPT-5.6 TerraOpenAI · 1.05M context———71.9▲0.1
15DeepSeek V4 ProDeepSeek · 1.049M context——71.8▼0.1
16GLM-5.2Z.ai · 1.049M context—71.5
17DeepSeek V4.1 FlashDeepSeek · 1.049M context———71.0▼0.2
18Gemini 3.7 FlashGoogle · 1.049M context——70.9
19Muse Spark 1.2Meta · 1.049M context—70.8
20Claude Sonnet 5Anthropic · 1M context—69.3▼0.1
21Gemini 3.1 ProGoogle · 1.049M context—69.2▼0.1
22Grok 4.6xAI · 500K context—68.9▼0.1
23InklingThinking Machines · 1.049M context—60.8▲0.5
24Kimi K2.7 CodeMoonshot AI · 262K context———59.5▲0.3
25Mistral Medium 3.5Mistral AI · 262K context——47.8▼0.1
26Seed 2.1 ProProvisionalByteDance Seed · 256K context—————79.6
27Hy4 previewProvisionalTencent · 1.049M context—————79.0
28Sakana NamazuProvisionalSakana AI · 262K context—————77.1
29Mercury 2.5ProvisionalInception · 260K context—————24.7
capability score 0–100 — tap a cell for the benchmarks and sources behind it▢ best in column▲▼ change vs. previous refreshNEW entered this week
Compare

Two-model comparison engine

Pick a model on each side. The stronger score per criterion turns green; the radar shows the shape of each model at a glance. Tap ⓘ for the benchmarks and sources behind a number.

Claude Mythos PreviewAnthropic
91.5thinQit Index
Blended price—
Output speed—
Context1M
CodingReasoningMathMultimodalAgenticHuman pref.
91Coding85
92Reasoning84
—Math—
—Multimodal—
—Agentic72
—Human pref.—
Claude Mythos Preview leads 2–0 on criteria · Index gap 14.9 pts. Claude Mythos Preview is the stronger all-rounder.
Claude Opus 5.5Anthropic
76.6thinQit Index
Blended price$8.00 / 1M
Output speed84 tok/s
Context1M
Local models

Will it run?

Pick your machine. We estimate memory fit and tokens per second for every open-weight model we track, at the best quantisation that fits.

Fast memory for weights48 GB
Memory bandwidth546 GB/s
Bytes/token budget @ 15 tok/s23.7 GB
18Run smoothly
0Run, but slowly
7Not realistic

Runs smoothly ≥ 15 tok/s on M4 Max

Muse Glimmer-30BMeta · 30B · Index 65.2Q4_K_Mrecommended20.6 GB17 tok/s
Qwen3.8 27BAlibaba · Qwen · 27.8B · 27B active · Index 61.9Q5_K_Mrecommended22.3 GB16 tok/s
Qwen3.6 27BAlibaba · Qwen · 27B · Index 61.2Q5_K_Mrecommended21.7 GB16 tok/s
Qwen3.6 35B-A3BAlibaba · Qwen · 35B · 3B active · Index 58.0Q8_0lossless40.5 GB73 tok/s
Qwen3.5 122B-A10BAlibaba · Qwen · 122B · 10B active · Index 56.2IQ2_XSheavy loss39.9 GB76 tok/s
Gemma 4 31BGoogle · 31B · Index 53.1Q4_K_Mrecommended21.2 GB17 tok/s
gpt-oss-120BOpenAI · 116.8B · 5.1B active · Index 49.9IQ2_XSheavy loss38.3 GB114 tok/s
Gemma 4 26B-A4BGoogle · 26B · 4B active · Index 46.1Q8_0lossless30.5 GB60 tok/s
gpt-oss-20BOpenAI · 20.9B · 3.6B active · Index 42.0Q8_0lossless24.8 GB64 tok/s
Qwen3.5 9BAlibaba · Qwen · 9B · Index 34.7Q8_0lossless11.5 GB31 tok/s
Gemma 4 12BGoogle · 12B · Index 32.6Q8_0lossless14.9 GB24 tok/s
Qwen3.5 4BAlibaba · Qwen · 4B · Index 30.3Q8_0lossless6.0 GB60 tok/s
Nemotron 3.5 Lightning 30B-A3BNVIDIA · 30B · 3B active · Index 28.3Q8_0lossless35.0 GB73 tok/s
Ministral 3 14BMistral AI · 14B · Index 22.9Q8_0lossless17.1 GB21 tok/s
Granite 4.2 8BIBM · 8.8B · Index 17.8Q8_0lossless11.3 GB31 tok/s
Llama 4 ScoutMeta · 108.6B · 17B active · Index 12.5IQ2_XSheavy loss35.7 GB52 tok/s
Gemma 4 E4BGoogle · 4B · Index 10.5Q8_0lossless6.0 GB60 tok/s
Llama 3.3 70BMeta · 70.6B · 70B active · Index 6.0IQ2_XSheavy loss23.7 GB15 tok/s

Runs, but slowly 3.75–15 tok/s — fine for background jobs

Nothing in this band

Not realistic below 3.75 tok/s or does not fit in memory

GLM-5.3Z.ai · 753.3B · ≈40B active · Index 72.4—481 GBdoes not fit— tok/s
DeepSeek V4 FlashDeepSeek · 304.2B · ≈24B active · Index 70.9—195 GBdoes not fit— tok/s
GLM-5.3-FlashZ.ai · 321.3B · ≈32B active · Index 69.8—206 GBdoes not fit— tok/s
Qwen3.8 2.4T-A95BAlibaba · Qwen · 2.4T · 95B active · Index 67.9—1529 GBdoes not fit— tok/s
Qwen3.8-Flash-NextAlibaba · Qwen · 177B · ≈12B active · Index 67.0—114 GBdoes not fit— tok/s
Qwen3.5 397B-A17BAlibaba · Qwen · 397B · 17B active · Index 59.7—254 GBdoes not fit— tok/s
Hy4 previewTencent · 780B · ≈40B active · Index 79.0—498 GBdoes not fit— tok/s
Insights

How the leaderboards agree with each other

Spearman rank correlation between sources over the models they share. Where two boards disagree, they measure something different — that disagreement is what earns a capability its weight in the Index. Cells need at least four shared models.

Artificial AnalysisLMArenaLLM-StatsLiveBenchSWE-benchScale SEALthinQit IndexArtificial Analysis1.000.640.930.50·0.260.93LMArena0.641.000.720.53·-0.400.85LLM-Stats0.930.721.000.61·0.140.94LiveBench0.500.530.611.00·0.100.77SWE-bench····1.00··Scale SEAL0.26-0.400.140.10·1.000.26thinQit Index0.930.850.940.77·0.261.00

Index v1.0 weights

25 Sept 2026
Reviewed against the correlation matrix at each refresh. A weight change of ≥2 points bumps the version and publishes a methodology note.
Coding25%
Reasoning20%
Agentic15%
Human pref.15%
Math10%
Multimodal10%
Speed & price5%
Across the AI landscape

One industry. Many different capabilities.

A review of the major AI capability areas, from reasoning and software delivery to creative media, voice and deployment. These are separate evaluations: their scores cannot be compared across rows or added to the thinQit Index.

Editorial review

The Index below refreshes automatically. This broader review is a dated editorial snapshot of public results and announcements. Source dates are shown where published; other links were checked on the review date.

01Language

Reasoning, math & science

Effort settings are part of the result

Claude Opus 5.5, released on 22 September, moved to the top of the Artificial Analysis Intelligence Index at 58 (max effort with fallback), five points clear of GPT-6 Astra and Claude Fable 5.1 at 53 and seven ahead of Claude Opus 5 at 51. Effort settings remain part of the result: Artificial Analysis publishes separate entries for the max, xhigh, high, medium and low settings, and the lower settings score well below the headline number.

How to read it Use reasoning and science evaluations alongside the aggregate. A rounded composite does not establish equal performance on every task.

02Code

Coding & software engineering

Track the task suite and the agent together

Anthropic reports Terminal-Bench 4.0 at 66.4% for Claude Opus 5.5, against 55.8% for Claude Fable 5.1 and 52.3% for Claude Opus 5. Artificial Analysis measures the same suite at 59.6% at max effort, a different harness and a different run, so the two numbers sit side by side rather than on one scale. Terminal-Bench 2.1 repairs 28 of the 89 tasks in version 2.0, Harbor-Index adds a broader agent evaluation, and SWE-bench remains a separate view of repository issue resolution.

How to read it Keep benchmark version, agent harness, model, resource limits and number of attempts attached to a score. A 2.0 result and a 2.1 result are different measurements.

03Agents

Agents & tool use

Real task completion gets its own view

Arena added GPT 6 Astra (Max) to Agent Arena on 8 September. Its latest board places Claude Fable 5.1 (Max) first and GPT 6 Astra (Max) second in the displayed order, with overlapping rank ranges of 1–4 and 1–5 respectively.

How to read it Read confirmed success, steerability, tool hallucination and cost per task together. These observational session metrics are not a controlled pass rate for your own workflows.

04Multimodal

Vision & document understanding

Reading an image differs from reading a document

Claude Fable 5 leads the displayed Vision Arena order at 1313 ±8. Claude Opus 5 (high) leads the separate Document Arena at 1520 ±15. The boards show different update dates and overlapping rank ranges.

How to read it Compare OCR, diagrams and long-document reasoning on matching tasks. Vision scores measure understanding; they do not measure image generation.

05Creative

Image generation & editing

New image leaders are still preliminary

Arena lists gpt-image-2.5-sunburst at 1421 ±13 for text-to-image and 1520 ±9 for single-image editing. Both are preliminary. The editing board also includes Microsoft’s MAI-Image-2.6, Meta’s muse-image, ByteDance’s Seedream and Google’s Nano Banana models.

How to read it Use the relevant category: text rendering, product design, photorealism or editing. Test whether an edit preserves the original subject and whether brand text is accurate.

06Creative

Video generation

Audio and input mode change the comparison

Artificial Analysis’s text-to-video board with audio lists Wan 3.0 at 1240 Elo, Gemini Omni Flash at 1239 and Minimax H3 Max (post-trained by fal) at 1235. Their confidence intervals overlap. Its recent additions include LTX-2.5 and MAGI-2 Preview.

How to read it Compare text-to-video, image-to-video and editing separately, with the same resolution, duration and audio setting. Preference scores and price per generated minute answer different questions.

07Audio

Speech generation

Voice choice is part of perceived quality

The current Artificial Analysis provider-voice table puts Cartesia Sonic 3.6 first at 1276 ±17 Elo. Inworld Realtime TTS-2 follows at 1245 ±18; Speechify Simba 3.2 and Qwen-Audio-3.0-TTS-Plus both show 1237 ±14.

How to read it This board uses each provider’s own voices. Audition your language, accent and speaking style; voice quality alone does not measure conversational latency or turn-taking.

08Audio

Transcription & real-time voice

Accuracy, response delay and turn-taking are separate

On AA-WER v2, MAI-Transcribe-2 records 2.0% word error rate, ElevenLabs Scribe v2 2.2%, and Mistral Voxtral Small 2.8%. The speech-to-speech evaluation separately measures audio reasoning, agentic performance and conversational dynamics.

How to read it Lower WER is better. Compare streaming and batch transcription separately. For conversations, inspect time to first audio, interruptions and pauses as well as transcription accuracy.

09Knowledge

Search, retrieval & embeddings

Grounded answers need more than a chat score

Search Arena’s latest displayed board lists GPT-5.6 Sol (xhigh) at 1257 ±7 and Claude Opus 4.6 Search at 1253 ±5, with overlapping rank ranges. MTEB evaluates embeddings and retrieval across languages and modalities as a separate layer.

How to read it Select retrieval tasks for your language and domain. Check citation correctness on your own documents; a preference ranking is not a factuality guarantee.

10Deployment

Open weights & local deployment

Small local models and frontier weights serve different needs

Artificial Analysis’s open-weight leaders include GLM-5.3 (max), Kimi K3 (max) and GLM-5.3-Flash. IBM’s Granite 4.2 family adds 3B, 8B and 30B options; its 8B model card provides an Apache-2.0 licence and reasoning/tool-use guidance.

How to read it Open weights do not imply consumer-hardware fit. Check total parameter memory, quantisation, context cache and licence; the hardware tool below provides estimates.

11Deployment

Cost, speed & new architectures

Evaluate the cost of completing the task

Anthropic priced Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens on 22 September, 20% below Claude Opus 5, with cache reads at $0.20. Artificial Analysis places four of its five effort settings on its intelligence versus cost per task frontier, and its blended rate at a 7:2:1 cache, input and output mix is $2.94 per million tokens. Inception's Mercury 2.5 reports 1,107 tokens/s and a 260K context window; those are vendor-reported figures, not thinQit measurements.

How to read it Compare time to first token, end-to-end completion, retries and cost per task. Keep launch discounts separate from standard prices and provider claims separate from independent results.

12Trust

Safety, robustness & hallucinations

Capability and safety need separate evidence

IBM’s Granite Guardian 4.1 adds configurable judging criteria and reports evaluations for RAG groundedness and tool-call hallucinations. HELM Safety offers a separate safety evaluation framework; its authors explicitly caution that it cannot certify a model as safe.

How to read it Inspect jailbreak resistance, false refusals, groundedness and tool permissions in the intended workflow. Vendor safety results and aggregate capability scores do not establish deployment readiness.

Image generation in practice Two original thinQit-style examples

Generated illustrations can follow an established editorial style or explore a more visual direction. These samples demonstrate composition, typography and style consistency; they are not scored benchmark results.

thinQit cover reading AI capabilities. The full picture., with a violet diagram connecting reasoning, code, image, video, voice and agents.
Editorial cover Dark grid, bold typography and a focused capability diagram.
thinQit illustration reading One brief. Many possibilities., with a glass prism connecting a puzzle, code, an image, a filmstrip, an audio wave and task nodes.
Illustrative direction The same palette expressed through glass, light and depth.
Sources

What we read every day

Each source is fetched by a scheduled job every morning. A red dot means the last read failed and its values are carried over from the previous successful run.

Benchmark news

What changed on the leaderboards

Written automatically when the daily refresh detects a new model, a rank change or a weight update — and published next to the regular blog.

All resources · blog & guides
Hardware guides

What it takes to run models locally

Regenerated from the local leaderboard and the Will it run? estimator, so the tables never drift from the tool.