Model benchmarks

mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF local LLM performance

As of August 2026, mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF runs at up to 193.7 tok/s for local inference (best of 2 community benchmark runs across 1 GPU).

vLLM
ShareRedditX

Model size

35B

Peak speed

193.7 tok/s

Average speed

190.9 tok/s

Min memory

24.2 GB

Max context

21,235 tokens

Best quality

89.3

Benchmark runs

2

GPUs tested

1

Performance by hardware and tool

Every hardware/tool/quantization combination mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
NVIDIA GeForce RTX 5090vLLM193.7 tok/s190.9 tok/s24.2 GB21,235 tokens82.12

Frequently asked questions

How fast is mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF for local inference?
Across 2 community benchmark runs, mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF reaches up to 193.7 tok/s and averages 190.9 tok/s, with the fastest results on NVIDIA GeForce RTX 5090.
How much memory does mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF need?
The leanest observed configuration used about 24.2 GB of memory.
Which tools have been used to run mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF?
Benchmarks were submitted using vLLM. Results are community-contributed and updated as new runs arrive.