Benchmark result
unsloth/Qwen3.5-2B-MTP-GGUF on NVIDIA GeForce RTX 4070 — 157.6 tok/s
Measured with Unsloth Studio on September 28, 2026.
What the model built
Open full screen →The coding scenario asks for a playable game in a single HTML file. This is exactly what the model returned, unedited.
How this run compares
7th fastest of 8 runs of this model on NVIDIA GeForce RTX 4070 · median 168.0 tok/s.
sort
- 1168.6tok/s50.1Unsloth Studio · UD-Q6_K_XL · KV f16
- 2168.5tok/s51.0Unsloth Studio · UD-Q6_K_XL · KV f16
- 3168.4tok/s36.8Unsloth Studio · UD-Q6_K_XL · KV f16
- 4168.1tok/s38.7Unsloth Studio · UD-Q6_K_XL · KV f16
- 5167.9tok/s37.0Unsloth Studio · UD-Q6_K_XL · KV f16
- 6167.9tok/s37.5Unsloth Studio · UD-Q6_K_XL · KV f16
- 7157.6tok/s49.1Unsloth Studio · UD-Q6_K_XL · KV f16this run
- 8141.4tok/s42.6Unsloth Studio · UD-Q6_K_XL · KV f16
unsloth/Qwen3.5-2B-MTP-GGUF on other hardware
Reproduce this run
toolUnsloth Studio
modelunsloth/Qwen3.5-2B-MTP-GGUF (UD-Q6_K_XL)
context65,536
kv cachef16
thinkingon
samplingtemperature 0.7 · top_p 0.8 · top_k 20 · min_p 0 · presence_penalty 1.5
clientv0.4.82+98
Set those in Unsloth Studio, then:
llm-benchmark benchmark --model "unsloth/Qwen3.5-2B-MTP-GGUF" --tool "Unsloth Studio"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.