Benchmark result
Qwen3.8-27B-oQ8e-mtp on Apple M5 Max — 39.7 tok/s
Measured with oMLX 0.31.4.dev132+g94cdcae13 on October 2, 2026.
What the model built
Open full screen →The coding scenario asks for a playable game in a single HTML file. This is exactly what the model returned, unedited; the frame only adds a line that tells this page how tall it is.
How this run compares
1st fastest of 33 runs of this model on Apple M5 Max · median 33.4 tok/s.
sort
- 139.7tok/s82.6oMLX · oQ8e · KV nonethis run
- 239.4tok/s78.7oMLX · KV none
- 339.1tok/s75.7oMLX
- 439.0tok/s81.2oMLX · KV none
- 538.6tok/s72.0oMLX
- 638.5tok/s76.0oMLX
- 737.3tok/s88.3oMLX · KV none
- 836.7tok/s70.0oMLX
- 935.6tok/s85.6oMLX · oQ8e · KV none
- 1035.4tok/s69.2oMLX
- 1135.0tok/s73.1oMLX
- 1234.4tok/s85.8oMLX
- 1333.9tok/s82.5oMLX
- 1433.7tok/s85.6oMLX
- 1533.7tok/s87.3oMLX · oQ8e · KV none
- 1633.5tok/s85.9oMLX
- 1733.4tok/s87.6oMLX · KV none
- 1833.2tok/s84.0oMLX
- 1932.9tok/s85.7oMLX · oQ8e · KV none
- 2032.7tok/s83.0oMLX
- 2132.4tok/s89.0oMLX
- 2232.3tok/s84.8oMLX
- 2332.2tok/s88.5oMLX
- 2432.1tok/s83.7oMLX · KV none
- 2532.1tok/s81.8oMLX · KV none
- 2631.8tok/s86.2oMLX
- 2730.8tok/s84.7oMLX
- 2830.0tok/s83.9oMLX
- 2929.3tok/s87.7oMLX · KV none
- 3027.5tok/s86.3oMLX · KV tq4
- 3127.3tok/s87.0oMLX
- 3227.1tok/s68.7oMLX
- 3320.4tok/s82.4oMLX
Reproduce this run
tooloMLX 0.31.4.dev132+g94cdcae13
modelQwen3.8-27B-oQ8e-mtp (oQ8e)
context262,144
kv cachenone
thinkingon
effortxhigh
think limitnone
samplingtemperature 1 · top_p 0.95 · top_k 20 · repetition_penalty 1
clientv0.4.87+102
Set those in oMLX, then:
llm-benchmark benchmark --model "Qwen3.8-27B-oQ8e-mtp" --tool "oMLX"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.