Benchmark result
qwen3:8b-q4_K_M on Apple M3 — 12.6 tok/s
Measured with Ollama 0.32.14 on August 19, 2026.
Reproduce this run
toolOllama 0.32.14
modelqwen3:8b-q4_K_M (Q4_K_M)
context8,192
clientv0.4.36+97
Set those in Ollama, then:
llm-benchmark benchmark --model "qwen3:8b-q4_K_M" --tool "Ollama"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.