A Mac can hold a 200B-parameter model in memory that no consumer GPU can — and then generate at a fraction of the speed. Choosing the right chip is about matching bandwidth to the models you care about.
Bandwidth, not cores
On Apple Silicon the CPU, GPU and Neural Engine share one memory pool, so 'VRAM' is simply most of your RAM (macOS lets you use roughly 75 % for models by default). What separates the chips is memory bandwidth, and it scales with the tier, not the generation:
| Chip | Bandwidth | Memory options |
|---|---|---|
| M4 | 120 GB/s | 16 GB / 24 GB / 32 GB |
| M4 Pro | 273 GB/s | 24 GB / 48 GB / 64 GB |
| M4 Max | 546 GB/s | 36 GB / 48 GB / 64 GB / 128 GB |
| M3 Ultra | 819 GB/s | 96 GB / 256 GB / 512 GB |
| M3 Max | 400 GB/s | 36 GB / 48 GB / 64 GB / 128 GB |
| M2 Ultra | 800 GB/s | 64 GB / 128 GB / 192 GB |
| M2 Max | 400 GB/s | 32 GB / 64 GB / 96 GB |
| M1 Max | 400 GB/s | 32 GB / 64 GB |
| M1 / M2 / M3 (base) | 100 GB/s | 8 GB / 16 GB / 24 GB |
An M4 Max (546 GB/s) generates roughly twice as fast as an M4 Pro (273 GB/s) on the same model, and an M3 Ultra (819 GB/s) another 1.5× on top. Base M-series chips (around 100–120 GB/s) are fine for 4–9B models and MoE models with a few billion active parameters, and painful for anything larger.
The table
| Model | Size | M4 · 24 GB | M4 Pro · 48 GB | M4 Max · 128 GB | M3 Ultra · 256 GB |
|---|---|---|---|---|---|
| GLM-5.3 | 753.3B · ≈40B active | — | — | — | — |
| DeepSeek V4 Flash | 304.2B · ≈24B active | — | — | — | Q3_K_M · 39 tok/s |
| GLM-5.3-Flash | 321.3B · ≈32B active | — | — | — | Q3_K_M · 30 tok/s |
| Qwen3.8 2.4T-A95B | 2.4T · 95B active | — | — | — | — |
| Qwen3.8-Flash-Next | 177B · ≈12B active | — | — | Q3_K_M · 46 tok/s | Q6_K · 45 tok/s |
| Muse Glimmer-30B | 30B | IQ2_XS · 7 tok/s | IQ2_XS · 16 tok/s | Q4_K_M · 17 tok/s | Q8_0 · 15 tok/s |
| Qwen3.8 27B | 27.8B · 27B active | IQ2_XS · 8 tok/s | IQ2_XS · 18 tok/s | Q5_K_M · 16 tok/s | Q8_0 · 17 tok/s |
| Qwen3.6 27B | 27B | IQ2_XS · 8 tok/s | IQ2_XS · 18 tok/s | Q5_K_M · 16 tok/s | Q8_0 · 17 tok/s |
| Qwen3.5 397B-A17B | 397B · 17B active | — | — | — | IQ2_XS · 78 tok/s |
| Qwen3.6 35B-A3B | 35B · 3B active | IQ2_XS · 32 tok/s | Q6_K · 43 tok/s | Q8_0 · 73 tok/s | Q8_0 · 110 tok/s |
| Qwen3.5 122B-A10B | 122B · 10B active | — | — | Q5_K_M · 40 tok/s | Q8_0 · 42 tok/s |
| Gemma 4 31B | 31B | IQ2_XS · 7 tok/s | IQ2_XS · 16 tok/s | Q4_K_M · 17 tok/s | Q6_K · 19 tok/s |
| gpt-oss-120B | 116.8B · 5.1B active | — | — | Q5_K_M · 67 tok/s | Q8_0 · 74 tok/s |
| Gemma 4 26B-A4B | 26B · 4B active | Q3_K_M · 22 tok/s | Q8_0 · 30 tok/s | Q8_0 · 60 tok/s | Q8_0 · 89 tok/s |
| gpt-oss-20B | 20.9B · 3.6B active | Q5_K_M · 19 tok/s | Q8_0 · 32 tok/s | Q8_0 · 64 tok/s | Q8_0 · 97 tok/s |
| Qwen3.5 9B | 9B | IQ2_XS · 18 tok/s | Q8_0 · 15 tok/s | Q8_0 · 31 tok/s | Q8_0 · 46 tok/s |
| Gemma 4 12B | 12B | Q4_K_M · 9 tok/s | Q5_K_M · 17 tok/s | Q8_0 · 24 tok/s | Q8_0 · 36 tok/s |
| Qwen3.5 4B | 4B | Q6_K · 16 tok/s | Q8_0 · 30 tok/s | Q8_0 · 60 tok/s | Q8_0 · 89 tok/s |
| Nemotron 3.5 Lightning 30B-A3B | 30B · 3B active | Q3_K_M · 26 tok/s | Q8_0 · 37 tok/s | Q8_0 · 73 tok/s | Q8_0 · 110 tok/s |
| Ministral 3 14B | 14B | Q4_K_M · 7 tok/s | Q4_K_M · 17 tok/s | Q8_0 · 21 tok/s | Q8_0 · 31 tok/s |
| Granite 4.2 8B | 8.8B | IQ2_XS · 18 tok/s | Q8_0 · 16 tok/s | Q8_0 · 31 tok/s | Q8_0 · 47 tok/s |
| Llama 4 Scout | 108.6B · 17B active | — | IQ2_XS · 26 tok/s | Q6_K · 22 tok/s | Q8_0 · 26 tok/s |
| Gemma 4 E4B | 4B | Q6_K · 16 tok/s | Q8_0 · 30 tok/s | Q8_0 · 60 tok/s | Q8_0 · 89 tok/s |
| Llama 3.3 70B | 70.6B · 70B active | — | IQ2_XS · 8 tok/s | IQ2_XS · 15 tok/s | IQ2_XS · 23 tok/s |
| Hy4 preview | 780B · ≈40B active | — | — | — | — |
Use MLX for the best speeds on Apple Silicon; llama.cpp with Metal is close behind and has the broadest quant support. Check your own configuration in the Will it run? tool.
Frequently asked questions
Is an M4 Max with 128 GB better than an RTX 5090 for local models?
Different trade-off. The 5090 is about 3× faster on anything that fits in its 32 GB; the Mac can load models four times larger. If you want 100B-class MoE models locally, the Mac wins; for 30B and below, the GPU wins on speed.
Why does macOS not let me use all my RAM for a model?
The system reserves memory for itself and other apps. The default GPU allocation limit is around 75 % of unified memory; it can be raised with sysctl, at the cost of stability when other apps need memory.
Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.
