Model benchmarks

Qwen3.8-9B-Q8_0:latest local LLM performance

As of August 2026, Qwen3.8-9B-Q8_0:latest runs at up to 76.5 tok/s for local inference (best of 2 community benchmark runs across 1 GPU).

OllamaQ8_0
ShareRedditX

Model size

9.2B

Peak speed

76.5 tok/s

Average speed

75.0 tok/s

Min memory

10.2 GB

Max context

7,505 tokens

Best quality

91.6

Benchmark runs

2

GPUs tested

1

Performance by hardware and tool

Every hardware/tool/quantization combination Qwen3.8-9B-Q8_0:latest has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
AMD Radeon RX 7900 XTXOllamaQ8_076.5 tok/s75.0 tok/s10.2 GB7,505 tokens72.82

Frequently asked questions

How fast is Qwen3.8-9B-Q8_0:latest for local inference?
Across 2 community benchmark runs, Qwen3.8-9B-Q8_0:latest reaches up to 76.5 tok/s and averages 75.0 tok/s, with the fastest results on AMD Radeon RX 7900 XTX.
How much memory does Qwen3.8-9B-Q8_0:latest need?
The leanest observed configuration used about 10.2 GB of memory (quantizations tested: Q8_0).
Which tools have been used to run Qwen3.8-9B-Q8_0:latest?
Benchmarks were submitted using Ollama. Results are community-contributed and updated as new runs arrive.