Hardware Benchmarks

Best Local LLM Performance for Apple M5 Pro

Compare Apple M5 Pro benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

Apple M5 Pro
VRAM
64GB

Recommended Models

Models we recommend for Apple M5 Pro.

Fastest

mtplx-qwen38-27b-oq4e-fp16-mtp

27B

486.9 tok/s

Highest token generation speed at 486.9 tok/s.

Best quality

scottlowry--Qwen3.8-27B-oQ4e-mtp

27B

32.2 tok/sQuality 85.1

Highest completed quality score at 85.1.

Largest context

qwen3:30b-thinking

30.5B

73.7 tok/sContext 65,536 tokens

Largest context that fits in 64GB VRAM, at 65,536 tokens.

Frequently asked questions

Which models actually run best on the Apple M5 Pro, by task, from community benchmark data.

What is the best local model for coding on the Apple M5 Pro?
On the Apple M5 Pro, scottlowry--Qwen3.8-27B-oQ4e-mtp ranks first for coding in our benchmarks (83.9/100 at ~32.2 tok/s).
What is the best local model for agentic workflows on the Apple M5 Pro?
On the Apple M5 Pro, qwen3.6-35b-a3b-ud-mlx ranks first for agentic workflows in our benchmarks (88.6/100 at ~71.3 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for Apple M5 Pro.