The Index below refreshes automatically. This broader review is a dated editorial snapshot of public results and announcements. Source dates are shown where published; other links were checked on the review date.
01Language
Reasoning, math & science
Effort settings are part of the result
Claude Opus 5.5, released on 22 September, moved to the top of the Artificial Analysis Intelligence Index at 58 (max effort with fallback), five points clear of GPT-6 Astra and Claude Fable 5.1 at 53 and seven ahead of Claude Opus 5 at 51. Effort settings remain part of the result: Artificial Analysis publishes separate entries for the max, xhigh, high, medium and low settings, and the lower settings score well below the headline number.
How to read it Use reasoning and science evaluations alongside the aggregate. A rounded composite does not establish equal performance on every task.
02Code
Coding & software engineering
Track the task suite and the agent together
Anthropic reports Terminal-Bench 4.0 at 66.4% for Claude Opus 5.5, against 55.8% for Claude Fable 5.1 and 52.3% for Claude Opus 5. Artificial Analysis measures the same suite at 59.6% at max effort, a different harness and a different run, so the two numbers sit side by side rather than on one scale. Terminal-Bench 2.1 repairs 28 of the 89 tasks in version 2.0, Harbor-Index adds a broader agent evaluation, and SWE-bench remains a separate view of repository issue resolution.
How to read it Keep benchmark version, agent harness, model, resource limits and number of attempts attached to a score. A 2.0 result and a 2.1 result are different measurements.
03Agents
Agents & tool use
Real task completion gets its own view
Arena added GPT 6 Astra (Max) to Agent Arena on 8 September. Its latest board places Claude Fable 5.1 (Max) first and GPT 6 Astra (Max) second in the displayed order, with overlapping rank ranges of 1–4 and 1–5 respectively.
How to read it Read confirmed success, steerability, tool hallucination and cost per task together. These observational session metrics are not a controlled pass rate for your own workflows.
04Multimodal
Vision & document understanding
Reading an image differs from reading a document
Claude Fable 5 leads the displayed Vision Arena order at 1313 ±8. Claude Opus 5 (high) leads the separate Document Arena at 1520 ±15. The boards show different update dates and overlapping rank ranges.
How to read it Compare OCR, diagrams and long-document reasoning on matching tasks. Vision scores measure understanding; they do not measure image generation.
05Creative
Image generation & editing
New image leaders are still preliminary
Arena lists gpt-image-2.5-sunburst at 1421 ±13 for text-to-image and 1520 ±9 for single-image editing. Both are preliminary. The editing board also includes Microsoft’s MAI-Image-2.6, Meta’s muse-image, ByteDance’s Seedream and Google’s Nano Banana models.
How to read it Use the relevant category: text rendering, product design, photorealism or editing. Test whether an edit preserves the original subject and whether brand text is accurate.
06Creative
Video generation
Audio and input mode change the comparison
Artificial Analysis’s text-to-video board with audio lists Wan 3.0 at 1240 Elo, Gemini Omni Flash at 1239 and Minimax H3 Max (post-trained by fal) at 1235. Their confidence intervals overlap. Its recent additions include LTX-2.5 and MAGI-2 Preview.
How to read it Compare text-to-video, image-to-video and editing separately, with the same resolution, duration and audio setting. Preference scores and price per generated minute answer different questions.
07Audio
Speech generation
Voice choice is part of perceived quality
The current Artificial Analysis provider-voice table puts Cartesia Sonic 3.6 first at 1276 ±17 Elo. Inworld Realtime TTS-2 follows at 1245 ±18; Speechify Simba 3.2 and Qwen-Audio-3.0-TTS-Plus both show 1237 ±14.
How to read it This board uses each provider’s own voices. Audition your language, accent and speaking style; voice quality alone does not measure conversational latency or turn-taking.
08Audio
Transcription & real-time voice
Accuracy, response delay and turn-taking are separate
On AA-WER v2, MAI-Transcribe-2 records 2.0% word error rate, ElevenLabs Scribe v2 2.2%, and Mistral Voxtral Small 2.8%. The speech-to-speech evaluation separately measures audio reasoning, agentic performance and conversational dynamics.
How to read it Lower WER is better. Compare streaming and batch transcription separately. For conversations, inspect time to first audio, interruptions and pauses as well as transcription accuracy.
09Knowledge
Search, retrieval & embeddings
Grounded answers need more than a chat score
Search Arena’s latest displayed board lists GPT-5.6 Sol (xhigh) at 1257 ±7 and Claude Opus 4.6 Search at 1253 ±5, with overlapping rank ranges. MTEB evaluates embeddings and retrieval across languages and modalities as a separate layer.
How to read it Select retrieval tasks for your language and domain. Check citation correctness on your own documents; a preference ranking is not a factuality guarantee.
10Deployment
Open weights & local deployment
Small local models and frontier weights serve different needs
Artificial Analysis’s open-weight leaders include GLM-5.3 (max), Kimi K3 (max) and GLM-5.3-Flash. IBM’s Granite 4.2 family adds 3B, 8B and 30B options; its 8B model card provides an Apache-2.0 licence and reasoning/tool-use guidance.
How to read it Open weights do not imply consumer-hardware fit. Check total parameter memory, quantisation, context cache and licence; the hardware tool below provides estimates.
11Deployment
Cost, speed & new architectures
Evaluate the cost of completing the task
Anthropic priced Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens on 22 September, 20% below Claude Opus 5, with cache reads at $0.20. Artificial Analysis places four of its five effort settings on its intelligence versus cost per task frontier, and its blended rate at a 7:2:1 cache, input and output mix is $2.94 per million tokens. Inception's Mercury 2.5 reports 1,107 tokens/s and a 260K context window; those are vendor-reported figures, not thinQit measurements.
How to read it Compare time to first token, end-to-end completion, retries and cost per task. Keep launch discounts separate from standard prices and provider claims separate from independent results.
12Trust
Safety, robustness & hallucinations
Capability and safety need separate evidence
IBM’s Granite Guardian 4.1 adds configurable judging criteria and reports evaluations for RAG groundedness and tool-call hallucinations. HELM Safety offers a separate safety evaluation framework; its authors explicitly caution that it cannot certify a model as safe.
How to read it Inspect jailbreak resistance, false refusals, groundedness and tool permissions in the intended workflow. Vendor safety results and aggregate capability scores do not establish deployment readiness.