Quantisation is the reason a 70B model runs on a laptop at all. It trades a little precision in the weights for a lot of memory — and the trade is far better than it sounds.
What the names mean
| Level | Bits / weight | Memory vs FP16 | Quality |
|---|---|---|---|
| FP16 / BF16 | 16 | 100 % | Reference — what the lab evaluated |
| Q8_0 | 8.5 | 53 % | lossless |
| Q6_K | 6.6 | 41 % | lossless |
| Q5_K_M | 5.7 | 36 % | recommended |
| Q4_K_M | 4.85 | 30 % | recommended |
| Q3_K_M | 3.9 | 24 % | some loss |
| IQ2_XS | 2.4 | 15 % | heavy loss |
The K-quants (Q4_K_M, Q5_K_M, Q6_K) keep the most sensitive tensors — attention and output layers — at higher precision than the name suggests, which is why Q4_K_M loses so little. IQ-quants (IQ2_XS, IQ3_S) use importance matrices to squeeze further; they are the only way to run 200B+ models on consumer hardware, but reasoning-heavy benchmarks drop measurably.
Memory per model at each level
| Model | Params | Q8_0 | Q5_K_M | Q4_K_M | IQ2_XS |
|---|---|---|---|---|---|
| GLM-5.3 | 753.3B · ≈40B active | 842 GB | 565 GB | 481 GB | 239 GB |
| DeepSeek V4 Flash | 304.2B · ≈24B active | 341 GB | 229 GB | 195 GB | 97.3 GB |
| GLM-5.3-Flash | 321.3B · ≈32B active | 360 GB | 242 GB | 206 GB | 103 GB |
| Qwen3.8 2.4T-A95B | 2.4T · 95B active | 2679 GB | 1797 GB | 1529 GB | 758 GB |
| Qwen3.8-Flash-Next | 177B · ≈12B active | 199 GB | 134 GB | 114 GB | 57.3 GB |
| Muse Glimmer-30B | 30B | 35.0 GB | 23.9 GB | 20.6 GB | 11.0 GB |
| Qwen3.8 27B | 27.8B · 27B active | 32.5 GB | 22.3 GB | 19.2 GB | 10.3 GB |
| Qwen3.6 27B | 27B | 31.6 GB | 21.7 GB | 18.7 GB | 10.0 GB |
| Qwen3.5 397B-A17B | 397B · 17B active | 444 GB | 299 GB | 254 GB | 127 GB |
| Qwen3.6 35B-A3B | 35B · 3B active | 40.5 GB | 27.7 GB | 23.8 GB | 12.5 GB |
| Qwen3.5 122B-A10B | 122B · 10B active | 138 GB | 92.8 GB | 79.2 GB | 39.9 GB |
| Gemma 4 31B | 31B | 36.1 GB | 24.7 GB | 21.2 GB | 11.3 GB |
| gpt-oss-120B | 116.8B · 5.1B active | 132 GB | 88.9 GB | 75.9 GB | 38.3 GB |
| Gemma 4 26B-A4B | 26B · 4B active | 30.5 GB | 21.0 GB | 18.1 GB | 9.7 GB |
Which to pick
The Will it run? tool applies these rules automatically: it picks the highest-quality quant that fits your memory and still reaches your target speed, and labels the result so you know when it had to compromise. Try it with your machine.
Frequently asked questions
Does quantisation make a model faster?
Yes, roughly in proportion to the memory saved: decode speed is bound by how many bytes are read per token, so Q4 generates about twice as fast as Q8 on the same hardware.
Do MoE models quantise as well as dense ones?
Generally yes, and their memory savings matter more because total parameters are large. Keep the shared attention layers at higher precision (the K-quants do this by default).
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
