Hardware Benchmarks

Best Local LLM Performance for Apple M2 Ultra

Compare Apple M2 Ultra benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

Apple M2 Ultra
VRAM
128GB

Recommended Models

Models we recommend for Apple M2 Ultra.

Fastest & Best quality & Largest context

Qwen3.8-27B-oQ8e-fp16-mtp

27B

34.9 tok/sQuality 83.9Context 262,144 tokens

Highest token generation speed at 34.9 tok/s. Highest completed quality score at 83.9. Largest context that fits in 128GB VRAM, at 262,144 tokens.

Frequently asked questions

Which models actually run best on the Apple M2 Ultra, by task, from community benchmark data.

What is the best local model for coding on the Apple M2 Ultra?
On the Apple M2 Ultra, Qwen3.8-27B-oQ8e-fp16-mtp ranks first for coding in our benchmarks (79.0/100 at ~34.4 tok/s).
What is the best local model for agentic workflows on the Apple M2 Ultra?
On the Apple M2 Ultra, Qwen3.6-27B-oQ8-mtp ranks first for agentic workflows in our benchmarks (80.0/100 at ~32.8 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for Apple M2 Ultra.