Benchmark result
Ornith-1.5-35B-A3B-APEX-MTP-I-Compact — 23.0 tok/s
Measured with vLLM on September 23, 2026.
What the model built
Open full screen →The coding scenario asks for a playable game in a single HTML file. This is exactly what the model returned, unedited.
How this run compares
Ornith-1.5-35B-A3B-APEX-MTP-I-Compact on other hardware
- 40.0tok/s85.6AMD Radeon RX 9050 / 9060 XT
- 23.0tok/s81.7Unknown hardwarethis run
Reproduce this run
toolvLLM
modelOrnith-1.5-35B-A3B-APEX-MTP-I-Compact
context98,304
thinkingon
clientv0.4.80+98
Set those in vLLM, then:
llm-benchmark benchmark --model "Ornith-1.5-35B-A3B-APEX-MTP-I-Compact" --tool "vLLM"The client prompts for context length and thinking mode, and for the KV cache dtype on Unsloth Studio. Sampling is left at the model default — the values above are what the tool reported using, not overrides the client sent.