Benchmark result
unsloth/Qwen3.8-27B-GGUF on 2× AMD Radeon RX 7900 XTX — 62.1 tok/s
Measured with Unsloth Studio on September 20, 2026.
What the model built
Open full screen →The coding scenario asks for a playable game in a single HTML file. This is exactly what the model returned, unedited; the frame only adds a line that tells this page how tall it is.
How this run compares
8th fastest of 21 runs of this model on 2× AMD Radeon RX 7900 XTX · Multi-GPU · median 60.0 tok/s.
sort
- 170.7tok/s83.1Unsloth Studio · UD-Q4_K_XL · KV f16
- 270.7tok/s81.5Unsloth Studio · UD-Q4_K_XL · KV f16
- 363.6tok/s81.9Unsloth Studio
- 462.7tok/s90.3Unsloth Studio · Q8_0 · KV f16
- 562.3tok/s86.4Unsloth Studio · Q8_0 · KV f16
- 662.3tok/s86.1Unsloth Studio · UD-Q6_K_XL · KV f16
- 762.3tok/s81.9Unsloth Studio · UD-Q4_K_XL · KV f16
- 862.1tok/s79.8Unsloth Studiothis run
- 961.2tok/s87.7Unsloth Studio
- 1061.0tok/s84.4Unsloth Studio
- 1160.0tok/s82.4Unsloth Studio
- 1259.8tok/s84.9Unsloth Studio
- 1358.0tok/s85.2Unsloth Studio
- 1451.4tok/s88.0Unsloth Studio
- 1543.1tok/s84.8Unsloth Studio · Q8_0 · KV q8_0
- 1640.8tok/s84.3Unsloth Studio
- 1740.5tok/s81.5Unsloth Studio
- 1839.4tok/s86.9Unsloth Studio · Q8_0 · KV q8_0
- 1938.8tok/s83.0Unsloth Studio · Q8_0 · KV q8_0
- 2038.3tok/s86.4Unsloth Studio · Q8_0 · KV q8_0
- 2125.6tok/s84.0Unsloth Studio · Q8_0 · KV q8_0
unsloth/Qwen3.8-27B-GGUF on other hardware
Reproduce this run
toolUnsloth Studio
modelunsloth/Qwen3.8-27B-GGUF
context65,536
clientv0.4.71+98
Set those in Unsloth Studio, then:
llm-benchmark benchmark --model "unsloth/Qwen3.8-27B-GGUF" --tool "Unsloth Studio"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.