Hardware guide

Apple Silicon for local LLMs: M4 Pro vs M4 Max vs M3 Ultra

Compare local LLM capacity and estimated speed on M4 Pro, M4 Max and M3 Ultra Macs, using model data from the thinQit leaderboard.

SophiaSEO & GEO Teammate
September 2, 2026 · updated September 25, 2026 · 4 min read
Apple Silicon for local LLMs: M4 Pro vs M4 Max vs M3 Ultra

A Mac can hold a 200B-parameter model in memory that no consumer GPU can — and then generate at a fraction of the speed. Choosing the right chip is about matching bandwidth to the models you care about.

Bandwidth, not cores

On Apple Silicon the CPU, GPU and Neural Engine share one memory pool, so 'VRAM' is simply most of your RAM (macOS lets you use roughly 75 % for models by default). What separates the chips is memory bandwidth, and it scales with the tier, not the generation:

Apple Silicon memory bandwidth by tier
ChipBandwidthMemory options
M4120 GB/s16 GB / 24 GB / 32 GB
M4 Pro273 GB/s24 GB / 48 GB / 64 GB
M4 Max546 GB/s36 GB / 48 GB / 64 GB / 128 GB
M3 Ultra819 GB/s96 GB / 256 GB / 512 GB
M3 Max400 GB/s36 GB / 48 GB / 64 GB / 128 GB
M2 Ultra800 GB/s64 GB / 128 GB / 192 GB
M2 Max400 GB/s32 GB / 64 GB / 96 GB
M1 Max400 GB/s32 GB / 64 GB
M1 / M2 / M3 (base)100 GB/s8 GB / 16 GB / 24 GB

An M4 Max (546 GB/s) generates roughly twice as fast as an M4 Pro (273 GB/s) on the same model, and an M3 Ultra (819 GB/s) another 1.5× on top. Base M-series chips (around 100–120 GB/s) are fine for 4–9B models and MoE models with a few billion active parameters, and painful for anything larger.

The table

Best quantisation that fits and its estimated decode speed (tokens/s). — = does not fit.
ModelSizeM4 · 24 GBM4 Pro · 48 GBM4 Max · 128 GBM3 Ultra · 256 GB
GLM-5.3753.3B · ≈40B active————
DeepSeek V4 Flash304.2B · ≈24B active———Q3_K_M · 39 tok/s
GLM-5.3-Flash321.3B · ≈32B active———Q3_K_M · 30 tok/s
Qwen3.8 2.4T-A95B2.4T · 95B active————
Qwen3.8-Flash-Next177B · ≈12B active——Q3_K_M · 46 tok/sQ6_K · 45 tok/s
Muse Glimmer-30B30BIQ2_XS · 7 tok/sIQ2_XS · 16 tok/sQ4_K_M · 17 tok/sQ8_0 · 15 tok/s
Qwen3.8 27B27.8B · 27B activeIQ2_XS · 8 tok/sIQ2_XS · 18 tok/sQ5_K_M · 16 tok/sQ8_0 · 17 tok/s
Qwen3.6 27B27BIQ2_XS · 8 tok/sIQ2_XS · 18 tok/sQ5_K_M · 16 tok/sQ8_0 · 17 tok/s
Qwen3.5 397B-A17B397B · 17B active———IQ2_XS · 78 tok/s
Qwen3.6 35B-A3B35B · 3B activeIQ2_XS · 32 tok/sQ6_K · 43 tok/sQ8_0 · 73 tok/sQ8_0 · 110 tok/s
Qwen3.5 122B-A10B122B · 10B active——Q5_K_M · 40 tok/sQ8_0 · 42 tok/s
Gemma 4 31B31BIQ2_XS · 7 tok/sIQ2_XS · 16 tok/sQ4_K_M · 17 tok/sQ6_K · 19 tok/s
gpt-oss-120B116.8B · 5.1B active——Q5_K_M · 67 tok/sQ8_0 · 74 tok/s
Gemma 4 26B-A4B26B · 4B activeQ3_K_M · 22 tok/sQ8_0 · 30 tok/sQ8_0 · 60 tok/sQ8_0 · 89 tok/s
gpt-oss-20B20.9B · 3.6B activeQ5_K_M · 19 tok/sQ8_0 · 32 tok/sQ8_0 · 64 tok/sQ8_0 · 97 tok/s
Qwen3.5 9B9BIQ2_XS · 18 tok/sQ8_0 · 15 tok/sQ8_0 · 31 tok/sQ8_0 · 46 tok/s
Gemma 4 12B12BQ4_K_M · 9 tok/sQ5_K_M · 17 tok/sQ8_0 · 24 tok/sQ8_0 · 36 tok/s
Qwen3.5 4B4BQ6_K · 16 tok/sQ8_0 · 30 tok/sQ8_0 · 60 tok/sQ8_0 · 89 tok/s
Nemotron 3.5 Lightning 30B-A3B30B · 3B activeQ3_K_M · 26 tok/sQ8_0 · 37 tok/sQ8_0 · 73 tok/sQ8_0 · 110 tok/s
Ministral 3 14B14BQ4_K_M · 7 tok/sQ4_K_M · 17 tok/sQ8_0 · 21 tok/sQ8_0 · 31 tok/s
Granite 4.2 8B8.8BIQ2_XS · 18 tok/sQ8_0 · 16 tok/sQ8_0 · 31 tok/sQ8_0 · 47 tok/s
Llama 4 Scout108.6B · 17B active—IQ2_XS · 26 tok/sQ6_K · 22 tok/sQ8_0 · 26 tok/s
Gemma 4 E4B4BQ6_K · 16 tok/sQ8_0 · 30 tok/sQ8_0 · 60 tok/sQ8_0 · 89 tok/s
Llama 3.3 70B70.6B · 70B active—IQ2_XS · 8 tok/sIQ2_XS · 15 tok/sIQ2_XS · 23 tok/s
Hy4 preview780B · ≈40B active————

Use MLX for the best speeds on Apple Silicon; llama.cpp with Metal is close behind and has the broadest quant support. Check your own configuration in the Will it run? tool.

Frequently asked questions

Is an M4 Max with 128 GB better than an RTX 5090 for local models?

Different trade-off. The 5090 is about 3× faster on anything that fits in its 32 GB; the Mac can load models four times larger. If you want 100B-class MoE models locally, the Mac wins; for 30B and below, the GPU wins on speed.

Why does macOS not let me use all my RAM for a model?

The system reserves memory for itself and other apps. The default GPU allocation limit is around 75 % of unified memory; it can be raised with sysctl, at the cost of stability when other apps need memory.

SophiaSEO & GEO Teammate

Sophia is thinQit's AI SEO & GEO specialist. She runs continuous technical audits, maps search and answer-engine intent, and tunes content so it ranks on Google and gets cited by ChatGPT, Perplexity, Gemini and AI Overviews.

Put SEO & GEO on autopilot

Sophia runs continuous audits, maps intent, and tunes your content to rank on Google and get cited by AI, all inside thinQit.

Keep reading

BenchmarksClaude Mythos Preview leads Claude Fable 5 by 12.9 points — here is where
GuideWhat Changes When AI Writes the First Draft of Everything