Model benchmarks

qwen3.8-27b local LLM performance

As of August 2026, qwen3.8-27b runs at up to 51.8 tok/s for local inference (best of 2 community benchmark runs across 2 GPUs).

LM Studio
ShareRedditX

Model size

27B

Peak speed

51.8 tok/s

Average speed

32.5 tok/s

Min memory

16.5 GB

Max context

8,192 tokens

Best quality

93.0

Benchmark runs

2

GPUs tested

2

Performance by hardware and tool

Every hardware/tool/quantization combination qwen3.8-27b has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
AMD Radeon RX 7900 XTXLM Studio51.8 tok/s51.8 tok/s16.5 GB8,192 tokens51.51
Apple M5 ProLM Studio13.3 tok/s13.3 tok/s17.6 GB8,192 tokens44.61

Frequently asked questions

How fast is qwen3.8-27b for local inference?
Across 2 community benchmark runs, qwen3.8-27b reaches up to 51.8 tok/s and averages 32.5 tok/s, with the fastest results on AMD Radeon RX 7900 XTX.
How much memory does qwen3.8-27b need?
The leanest observed configuration used about 16.5 GB of memory.
Which tools have been used to run qwen3.8-27b?
Benchmarks were submitted using LM Studio. Results are community-contributed and updated as new runs arrive.