Benchmark result
unsloth/Qwen3.8-27B-GGUF on 2× AMD Radeon RX 7900 XTX — 38.8 tok/s
Measured with Unsloth Studio on September 28, 2026.
What the model built
Open full screen →The coding scenario asks for a playable game in a single HTML file. This is exactly what the model returned, unedited; the frame only adds a line that tells this page how tall it is.
How this run compares
19th fastest of 21 runs of this model on 2× AMD Radeon RX 7900 XTX · Multi-GPU · median 60.0 tok/s.
sort
- 170.7tok/s83.1Unsloth Studio · UD-Q4_K_XL · KV f16
- 270.7tok/s81.5Unsloth Studio · UD-Q4_K_XL · KV f16
- 363.6tok/s81.9Unsloth Studio
- 462.7tok/s90.3Unsloth Studio · Q8_0 · KV f16
- 562.3tok/s86.4Unsloth Studio · Q8_0 · KV f16
- 662.3tok/s86.1Unsloth Studio · UD-Q6_K_XL · KV f16
- 762.3tok/s81.9Unsloth Studio · UD-Q4_K_XL · KV f16
- 862.1tok/s79.8Unsloth Studio
- 961.2tok/s87.7Unsloth Studio
- 1061.0tok/s84.4Unsloth Studio
- 1160.0tok/s82.4Unsloth Studio
- 1259.8tok/s84.9Unsloth Studio
- 1358.0tok/s85.2Unsloth Studio
- 1451.4tok/s88.0Unsloth Studio
- 1543.1tok/s84.8Unsloth Studio · Q8_0 · KV q8_0
- 1640.8tok/s84.3Unsloth Studio
- 1740.5tok/s81.5Unsloth Studio
- 1839.4tok/s86.9Unsloth Studio · Q8_0 · KV q8_0
- 1938.8tok/s83.0Unsloth Studio · Q8_0 · KV q8_0this run
- 2038.3tok/s86.4Unsloth Studio · Q8_0 · KV q8_0
- 2125.6tok/s84.0Unsloth Studio · Q8_0 · KV q8_0
unsloth/Qwen3.8-27B-GGUF on other hardware
Reproduce this run
toolUnsloth Studio
modelunsloth/Qwen3.8-27B-GGUF (Q8_0)
context131,072
kv cacheq8_0 (assumed default)
effortxhigh (model default)
think limit8,192 tokens
samplingtemperature 0.7 · top_p 0.8 · top_k 20 · min_p 0 · presence_penalty 1.5
clientv0.4.82+98
Set those in Unsloth Studio, then:
llm-benchmark benchmark --model "unsloth/Qwen3.8-27B-GGUF" --tool "Unsloth Studio"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.