Fastest
qwen2.5-coder:7b-base-q4_K_M
7.6B
19.1 tok/s
Highest token generation speed at 19.1 tok/s.
Hardware Benchmarks
Compare Apple M4 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for Apple M4.
Fastest
7.6B
Highest token generation speed at 19.1 tok/s.
Best quality & Largest context
9.0B
Highest completed quality score at 57.4. Largest context that fits in 16GB VRAM, at 8,192 tokens.
Which models actually run best on the Apple M4, by task, from community benchmark data.
Completed quality-ranked benchmark results for Apple M4.