Every lab publishes the benchmarks it wins. Every leaderboard measures something slightly different. The thinQit Index is our attempt at a single, honest, daily answer to 'which model is best right now — and best at what?'
Why another score
We build products on these models every day, and the question we actually get from clients is not 'what is GPT's GPQA score' but 'should we use Claude, GPT or Gemini for this?' Answering that means reading eight leaderboards, mentally normalising Elo ratings against percentages, and remembering which numbers are a month old. The Index does that reading for us, in code, every morning.
How it is computed
- Collect. A GitHub Actions job reads the public leaderboards daily: Artificial Analysis, LMArena, LLM-Stats (SWE-bench Verified, GPQA Diamond, Humanity's Last Exam, AIME, LiveCodeBench, MMMU, Terminal-Bench), LiveBench, SWE-bench, Scale SEAL, plus Hugging Face and OpenRouter for parameter counts, licences, context windows and prices.
- Normalise. Each benchmark is mapped to 0–100 with fixed anchors (for example SWE-bench Verified 30 % → 0, 100 % → 100; LMArena 1250 → 0, 1550 → 100). Fixed anchors keep yesterday's scores comparable with today's, which a min–max over the current field would not.
- Roll up. Benchmarks are averaged inside their capability: coding, reasoning, math, multimodal, agentic, human preference, and speed & price.
- Weight. The Index is a weighted mean over the capabilities a model actually has results for — coding 25 %, reasoning 20 %, agentic 15 %, human preference 15 %, math 10 %, multimodal 10 %, speed & price 5 %. Missing capabilities are excluded, never imputed. A model needs at least two substantive capabilities to be ranked; until then it shows as provisional.
When the leaderboards disagree
We publish the rank correlation between every pair of sources on the benchmarks page. Where two boards agree closely, they are measuring the same thing and one of them is adding little; where they diverge, a capability is doing real work. That matrix is what we use to review the weights, and any weight change of two points or more bumps the Index version and gets its own post here.
Local models and 'Will it run?'
The same Index covers open-weight models you can run yourself, from 4B phones-and-laptops models to trillion-parameter mixtures of experts. Because a score is meaningless if you cannot run the model, the page includes a Will it run? tool: pick your chip, memory and what 'smooth' means to you, and it sorts every local model into runs smoothly, runs slowly, or not realistic — with the quantisation it would pick and an estimated tokens per second. Explore it on the AI Benchmarks page.
Frequently asked questions
How often is the thinQit Index updated?
Daily. A scheduled job re-reads the public leaderboards, recomputes every score and publishes a news post only when a ranking, a model list or the methodology changed.
Can a model score high on the Index with only one benchmark?
No. A model needs results in at least two substantive capabilities to be ranked; with fewer it is shown as provisional and sorts below ranked models regardless of its score.
Where do the numbers come from?
Every cell on the page has an info button listing the exact benchmark, the source leaderboard, the value and the date it was observed. Nothing is estimated except the Will it run? speeds, which are labelled as estimates.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
