VRAM decides what you can load; memory bandwidth decides how fast it answers. This guide puts every open-weight model on the thinQit local leaderboard against the GPUs people actually buy, using the same estimator as the Will it run? tool.
The two numbers that matter
A model's weights have to sit somewhere the GPU can read them quickly. At Q4_K_M (about 4.85 bits per weight) a dense model needs roughly 0.6 GB per billion parameters, plus 1–2 GB for the KV cache. That is the first number: does it fit in VRAM? The second number is bandwidth: once the weights fit, each generated token requires the GPU to read every active weight once, so tokens/s ≈ bandwidth ÷ bytes read per token. An RTX 4090 moves about 1 TB/s; an RTX 4070 about half that.
The table
| Model | Size | RTX 4070 · 12 GB | RTX 5080 · 16 GB | RTX 4090 · 24 GB | RTX 5090 · 32 GB | 2× RTX 4090 · 48 GB |
|---|---|---|---|---|---|---|
| GLM-5.3 | 753.3B · ≈40B active | — | — | — | — | — |
| DeepSeek V4 Flash | 304.2B · ≈24B active | — | — | — | — | — |
| GLM-5.3-Flash | 321.3B · ≈32B active | — | — | — | — | — |
| Qwen3.8 2.4T-A95B | 2.4T · 95B active | — | — | — | — | — |
| Qwen3.8-Flash-Next | 177B · ≈12B active | — | — | — | — | — |
| Muse Glimmer-30B | 30B | IQ2_XS · 30 tok/s | IQ2_XS · 57 tok/s | Q4_K_M · 32 tok/s | Q6_K · 42 tok/s | Q8_0 · 19 tok/s |
| Qwen3.8 27B | 27.8B · 27B active | IQ2_XS · 33 tok/s | IQ2_XS · 62 tok/s | Q5_K_M · 30 tok/s | Q6_K · 47 tok/s | Q8_0 · 21 tok/s |
| Qwen3.6 27B | 27B | IQ2_XS · 33 tok/s | IQ2_XS · 62 tok/s | Q5_K_M · 30 tok/s | Q6_K · 47 tok/s | Q8_0 · 21 tok/s |
| Qwen3.5 397B-A17B | 397B · 17B active | — | — | — | — | — |
| Qwen3.6 35B-A3B | 35B · 3B active | Q4_K_M · 17 tok/s (RAM spill) | IQ2_XS · 255 tok/s | Q3_K_M · 216 tok/s | Q5_K_M · 311 tok/s | Q8_0 · 135 tok/s |
| Qwen3.5 122B-A10B | 122B · 10B active | — | — | — | — | IQ2_XS · 141 tok/s |
| Gemma 4 31B | 31B | Q4_K_M · 3 tok/s (RAM spill) | IQ2_XS · 55 tok/s | Q4_K_M · 31 tok/s | Q6_K · 41 tok/s | Q8_0 · 18 tok/s |
| gpt-oss-120B | 116.8B · 5.1B active | — | — | — | — | IQ2_XS · 211 tok/s |
| Gemma 4 26B-A4B | 26B · 4B active | IQ2_XS · 119 tok/s | Q3_K_M · 176 tok/s | Q5_K_M · 146 tok/s | Q8_0 · 195 tok/s | Q8_0 · 110 tok/s |
| gpt-oss-20B | 20.9B · 3.6B active | IQ2_XS · 124 tok/s | Q4_K_M · 165 tok/s | Q6_K · 142 tok/s | Q8_0 · 211 tok/s | Q8_0 · 119 tok/s |
| Qwen3.5 9B | 9B | Q6_K · 35 tok/s | Q8_0 · 54 tok/s | Q8_0 · 57 tok/s | Q8_0 · 101 tok/s | Q8_0 · 57 tok/s |
| Gemma 4 12B | 12B | Q5_K_M · 31 tok/s | Q8_0 · 42 tok/s | Q8_0 · 44 tok/s | Q8_0 · 78 tok/s | Q8_0 · 44 tok/s |
| Qwen3.5 4B | 4B | Q8_0 · 55 tok/s | Q8_0 · 105 tok/s | Q8_0 · 110 tok/s | Q8_0 · 195 tok/s | Q8_0 · 110 tok/s |
| Nemotron 3.5 Lightning 30B-A3B | 30B · 3B active | IQ2_XS · 134 tok/s | IQ2_XS · 255 tok/s | Q4_K_M · 192 tok/s | Q6_K · 284 tok/s | Q8_0 · 135 tok/s |
| Ministral 3 14B | 14B | Q4_K_M · 31 tok/s | Q6_K · 46 tok/s | Q8_0 · 38 tok/s | Q8_0 · 68 tok/s | Q8_0 · 38 tok/s |
| Granite 4.2 8B | 8.8B | Q6_K · 36 tok/s | Q8_0 · 55 tok/s | Q8_0 · 58 tok/s | Q8_0 · 103 tok/s | Q8_0 · 58 tok/s |
| Llama 4 Scout | 108.6B · 17B active | — | — | — | — | IQ2_XS · 96 tok/s |
| Gemma 4 E4B | 4B | Q8_0 · 55 tok/s | Q8_0 · 105 tok/s | Q8_0 · 110 tok/s | Q8_0 · 195 tok/s | Q8_0 · 110 tok/s |
| Llama 3.3 70B | 70.6B · 70B active | — | — | Q4_K_M · 1 tok/s (RAM spill) | IQ2_XS · 49 tok/s | Q3_K_M · 18 tok/s |
| Hy4 preview | 780B · ≈40B active | — | — | — | — | — |
Estimates assume a 32 GB DDR5 system (64 GB for the dual-4090 rig), a 4K context, and llama.cpp or vLLM with the whole model on the GPU where it fits. Real numbers move ±30 % with driver stack, prompt length and sampling settings. Try your own configuration in the Will it run? tool.
Which quantisation to pick
Q5_K_M and Q4_K_M are the workhorses: quality loss is small on benchmarks and imperceptible in most chat use. Q8_0 is near-lossless but doubles memory versus Q4; use it when the model is small enough that speed is not the constraint. Q3 and IQ2 quants exist so that a 100B+ model fits at all — expect noticeably weaker reasoning, and treat their Index scores as an upper bound.
Frequently asked questions
Does a faster GPU help if the model does not fit?
No. Once part of the model lives in system RAM, DDR bandwidth dominates and the GPU spends most of its time waiting. Fitting the model in VRAM matters more than GPU generation.
Are these numbers measured or estimated?
Estimated from memory bandwidth and bytes per token, calibrated against community llama.cpp and vLLM runs. They are meant to separate 'runs well', 'runs slowly' and 'does not run', not to replace a benchmark on your own machine.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
