Benchmark result
gpt-oss:20b on AMD Radeon RX 7900 XTX — 48.6 tok/s
Measured with Ollama 0.20.7 on April 16, 2026.
What the model built
Open full screen →The coding scenario asks for a playable game in a single HTML file. This is exactly what the model returned, unedited; the frame only adds a line that tells this page how tall it is.
How this run compares
20th fastest of 20 runs of this model on AMD Radeon RX 7900 XTX · median 69.9 tok/s.
sort
- 1111.1tok/s63.8LM Studio
- 2111.0tok/s74.6LM Studio
- 3110.8tok/s75.8LM Studio
- 4110.0tok/s64.6LM Studio
- 5109.6tok/s64.2LM Studio
- 6108.3tok/s62.5LM Studio
- 7108.3tok/s63.4LM Studio
- 8106.7tok/s67.8LM Studio
- 9106.2tok/s66.2LM Studio
- 1070.6tok/s72.6Ollama · MXFP4
- 1169.1tok/s75.4Ollama · MXFP4
- 1267.8tok/s75.9Ollama · MXFP4
- 1366.7tok/s74.5Ollama · MXFP4
- 1464.9tok/s71.6Ollama · MXFP4
- 1563.5tok/s77.9Ollama · MXFP4
- 1662.1tok/s73.4Ollama · MXFP4
- 1760.5tok/s73.7Ollama · MXFP4
- 1857.0tok/s75.0Ollama · MXFP4
- 1951.3tok/s73.8Ollama · MXFP4
- 2048.6tok/s64.5Ollama · MXFP4this run
Reproduce this run
toolOllama 0.20.7
modelgpt-oss:20b (MXFP4)
context4,096
clientv0.4.9+85
Set those in Ollama, then:
llm-benchmark benchmark --model "gpt-oss:20b" --tool "Ollama"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.