Hardware Benchmarks

Best Local LLM Performance for RX 9070 9070 Xt 9070 Gre

Compare RX 9070 9070 Xt 9070 Gre benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

AMD Radeon RX 9070/9070 XT/9070 GRE
VRAM
15GB

Recommended Models

Models we recommend for RX 9070 9070 Xt 9070 Gre.

Fastest

/models/Ornith-1.5-9B-Q8_0.gguf

9B

54.9 tok/s

Highest token generation speed at 54.9 tok/s.

Best quality

/models/Qwen3.8-27B-UD-IQ4_XS.gguf

27B

30.8 tok/sQuality 84.3

Highest completed quality score at 84.3.

Largest context

/models/gemma4-v2-Q6_K.gguf

Unknown

47.8 tok/sContext 65,536 tokens

Largest context that fits in 15GB VRAM, at 65,536 tokens.

Frequently asked questions

Which models actually run best on the RX 9070 9070 Xt 9070 Gre, by task, from community benchmark data.

What is the best local model for coding on the RX 9070 9070 Xt 9070 Gre?
On the RX 9070 9070 Xt 9070 Gre, /models/Qwen3.8-27B-UD-IQ4_XS.gguf ranks first for coding in our benchmarks (77.8/100 at ~30.8 tok/s).
What is the best local model for agentic workflows on the RX 9070 9070 Xt 9070 Gre?
On the RX 9070 9070 Xt 9070 Gre, /models/Qwen3.8-27B-UD-IQ4_XS.gguf ranks first for agentic workflows in our benchmarks (82.9/100 at ~30.8 tok/s). /models/gemma4-v2-Q6_K.gguf is faster (~47.8 tok/s) and still scores well (75.6/100) — a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for AMD Radeon RX 9070/9070 XT/9070 GRE.