Model benchmarks

Qwen3.8:27B/MTP@180K local LLM performance

As of August 2026, Qwen3.8:27B/MTP@180K runs at up to 53.0 tok/s for local inference (best of 2 community benchmark runs across 0 GPUs).

LM Studio
ShareRedditX

Model size

27B

Peak speed

53.0 tok/s

Average speed

47.6 tok/s

Min memory

13.2 GB

Max context

14,274 tokens

Best quality

93.3

Benchmark runs

2

GPUs tested

0

Performance by hardware and tool

Every hardware/tool/quantization combination Qwen3.8:27B/MTP@180K has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
CPU onlyLM Studio53.0 tok/s47.6 tok/s13.2 GB14,274 tokens57.42

Frequently asked questions

How fast is Qwen3.8:27B/MTP@180K for local inference?
Across 2 community benchmark runs, Qwen3.8:27B/MTP@180K reaches up to 53.0 tok/s and averages 47.6 tok/s.
How much memory does Qwen3.8:27B/MTP@180K need?
The leanest observed configuration used about 13.2 GB of memory.
Which tools have been used to run Qwen3.8:27B/MTP@180K?
Benchmarks were submitted using LM Studio. Results are community-contributed and updated as new runs arrive.