Hardware Benchmarks

Best Local LLM Performance for Apple M5 Ultra

Compare Apple M5 Ultra benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

Apple M5 Ultra
VRAM
256GB

Recommended Models

Models we recommend for Apple M5 Ultra.

Fastest

Tiel-Coder-35B-A3B-MLX-oQ6e-MTP

35B

214.4 tok/s

Highest token generation speed at 214.4 tok/s.

Best quality

ThinkingCap-Qwen3.8-27B-oQ4e-mtp

27B

93.6 tok/sQuality 88.8

Highest completed quality score at 88.8.

Frequently asked questions

Which models actually run best on the Apple M5 Ultra, by task, from community benchmark data.

What is the best local model for coding on the Apple M5 Ultra?
On the Apple M5 Ultra, ThinkingCap-Qwen3.8-27B-oQ4e-mtp ranks first for coding in our benchmarks (84.6/100 at ~93.6 tok/s).
What is the best local model for agentic workflows on the Apple M5 Ultra?
On the Apple M5 Ultra, ThinkingCap-Qwen3.8-27B-oQ4e-mtp ranks first for agentic workflows in our benchmarks (89.1/100 at ~93.6 tok/s). Qwen3.8-Flash-Next-oQ4e-mtp:coding-xhigh is faster (~130.9 tok/s) and still scores well (85.2/100) — a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for Apple M5 Ultra.