Hardware Benchmarks

Best Local LLM Performance for Instinct Mi60

Compare Instinct Mi60 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

Radeon Instinct MI60
VRAM
31GB

Recommended Models

Models we recommend for Instinct Mi60.

Fastest & Largest context

C:\Users\Choco\Downloads\gemma-4-26B-A4B-it-qat-UD-Q4_K_XL\gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf

26B

76.8 tok/sContext 65,536 tokens

Highest token generation speed at 76.8 tok/s. Largest context that fits in 31GB VRAM, at 65,536 tokens.

Best quality

Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf

35B

58.4 tok/sQuality 76.7

Highest completed quality score at 76.7.

Frequently asked questions

Which models actually run best on the Instinct Mi60, by task, from community benchmark data.

What is the best local model for coding on the Instinct Mi60?
On the Instinct Mi60, Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf ranks first for coding in our benchmarks (71.8/100 at ~58.4 tok/s).
What is the best local model for agentic workflows on the Instinct Mi60?
On the Instinct Mi60, Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf ranks first for agentic workflows in our benchmarks (78.8/100 at ~58.4 tok/s). C:\Users\Choco\Downloads\gemma-4-26B-A4B-it-qat-UD-Q4_K_XL\gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf is faster (~76.8 tok/s) and still scores well (76.5/100) — a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for Radeon Instinct MI60.