Fastest
qwen3.5:0.8b-mlx
852.88M
223.0 tok/s
Highest token generation speed at 223.0 tok/s.
Hardware Benchmarks
Compare Apple M3 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for Apple M3.
Fastest
852.88M
Highest token generation speed at 223.0 tok/s.
Best quality
20.9B
Highest completed quality score at 71.2.
Largest context
8.2B
Largest context that fits in 16GB VRAM, at 8,192 tokens.
Completed quality-ranked benchmark results for Apple M3.