Hardware Benchmarks

Best Local LLM Performance for Apple M4 Pro

Compare Apple M4 Pro benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

Apple M4 Pro
VRAM
48GB

Recommended Models

Models we recommend for Apple M4 Pro.

Fastest & Best quality

Ornith-1.5-35B-A3B-oQ4e-mtp

35B

69.4 tok/sQuality 86.2

Highest token generation speed at 69.4 tok/s. Highest completed quality score at 86.2.

Largest context

gemma4:26b-a4b-it-qat

25.2B

45.3 tok/sContext 524,288 tokens

Largest context that fits in 48GB VRAM, at 524,288 tokens.

Frequently asked questions

Which models actually run best on the Apple M4 Pro, by task, from community benchmark data.

What is the best local model for coding on the Apple M4 Pro?
On the Apple M4 Pro, Ornith-1.5-35B-A3B-oQ4e-mtp ranks first for coding in our benchmarks (79.0/100 at ~69.4 tok/s).
What is the best local model for agentic workflows on the Apple M4 Pro?
On the Apple M4 Pro, Ornith-1.5-35B-A3B-oQ4e-mtp ranks first for agentic workflows in our benchmarks (91.3/100 at ~69.4 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for Apple M4 Pro.