Fastest & Largest context
qwen3:8b-q4_K_M
8.2B
12.6 tok/sContext 8,192 tokens
Highest token generation speed at 12.6 tok/s. Largest context that fits in 16GB VRAM, at 8,192 tokens.
Hardware Benchmarks
Compare Apple M3 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for Apple M3.
Fastest & Largest context
8.2B
Highest token generation speed at 12.6 tok/s. Largest context that fits in 16GB VRAM, at 8,192 tokens.
Best quality
20.9B
Highest completed quality score at 74.3.
Which models actually run best on the Apple M3, by task, from community benchmark data.
Completed quality-ranked benchmark results for Apple M3.