Benchmark result
Qwen3.8-27B-oQ8e-mtp on Apple M5 Max — 30.0 tok/s
Measured with oMLX 0.31.3 on August 28, 2026.
What the model built
Open full screen →The coding scenario asks for a playable game in a single HTML file. This is exactly what the model returned, unedited.
How this run compares
27th fastest of 32 runs of this model on Apple M5 Max · median 33.3 tok/s.
sort
- 139.4tok/s78.7oMLX · KV none
- 239.1tok/s75.7oMLX
- 339.0tok/s81.2oMLX · KV none
- 438.6tok/s72.0oMLX
- 538.5tok/s76.0oMLX
- 637.3tok/s88.3oMLX · KV none
- 736.7tok/s70.0oMLX
- 835.6tok/s85.6oMLX · oQ8e · KV none
- 935.4tok/s69.2oMLX
- 1035.0tok/s73.1oMLX
- 1134.4tok/s85.8oMLX
- 1233.9tok/s82.5oMLX
- 1333.7tok/s85.6oMLX
- 1433.7tok/s87.3oMLX · oQ8e · KV none
- 1533.5tok/s85.9oMLX
- 1633.4tok/s87.6oMLX · KV none
- 1733.2tok/s84.0oMLX
- 1832.9tok/s85.7oMLX · oQ8e · KV none
- 1932.7tok/s83.0oMLX
- 2032.4tok/s89.0oMLX
- 2132.3tok/s84.8oMLX
- 2232.2tok/s88.5oMLX
- 2332.1tok/s83.7oMLX · KV none
- 2432.1tok/s81.8oMLX · KV none
- 2531.8tok/s86.2oMLX
- 2630.8tok/s84.7oMLX
- 2730.0tok/s83.9oMLXthis run
- 2829.3tok/s87.7oMLX · KV none
- 2927.5tok/s86.3oMLX · KV tq4
- 3027.3tok/s87.0oMLX
- 3127.1tok/s68.7oMLX
- 3220.4tok/s82.4oMLX
Reproduce this run
tooloMLX 0.31.3
modelQwen3.8-27B-oQ8e-mtp
context262,144
thinkingon
clientv0.4.44+97
Set those in oMLX, then:
llm-benchmark benchmark --model "Qwen3.8-27B-oQ8e-mtp" --tool "oMLX"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.