Model benchmarks

qwen3.8-27b local LLM performance

As of September 2026, qwen3.8-27b runs at up to 156.8 tok/s for local inference (best of 26 community benchmark runs across 12 GPUs).

LM StudioOllamallama.cppninfervLLMQ4_K_M
ShareRedditX

Model size

27B

Peak speed

156.8 tok/s

Average speed

50.3 tok/s

Avg PP

1365.2 tok/s

Min memory

16.3 GB

Max context

123,904 tokens

Avg output / run

49,550 tokens

Avg runtime / run

24m 14s

Avg quality

73.1

Benchmark runs

26

GPUs tested

12

Quality by task

Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 26 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.

TaskP5 (low)AvgP95 (high)
Overall44.773.188.0
Agent Workflow74.384.791.1
Code Generation1.651.482.9
Role Play & Narrative79.889.094.5
Research & Analysis0.367.490.3

Performance by hardware and tool

Every hardware/tool/quantization combination qwen3.8-27b has been benchmarked on, ranked by peak token generation speed. Last updated September 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
NVIDIA GeForce RTX 5090ninfer—156.8 tok/s150.0 tok/sn/a65,536 tokens68.02
NVIDIA GeForce RTX 5090vLLM—155.2 tok/s155.2 tok/sn/a65,536 tokens89.11
NVIDIA L4OllamaQ4_K_M109.3 tok/s109.3 tok/s16.7 GB65,536 tokens86.71
AMD Radeon RX 7900 XT/7900 XTX/7900Mllama.cpp—52.4 tok/s51.8 tok/sn/a65,536 tokens58.94
AMD Radeon RX 7900 XT/7900 XTX/7900MLM Studio—51.9 tok/s51.9 tok/sn/a65,536 tokens86.31
AMD Radeon RX 7900 XTXLM Studio—51.8 tok/s51.8 tok/sn/a65,536 tokens45.11
NVIDIA GeForce RTX 5070 TiLM Studio—46.9 tok/s42.8 tok/sn/a121,072 tokens87.92
AMD Radeon RX 7900 XTXOllamaQ4_K_M39.4 tok/s38.4 tok/s16.3 GB65,536 tokens82.24
NVIDIA GeForce RTX 4090LM Studio—39.0 tok/s39.0 tok/sn/a8,192 tokens66.71
AMD Radeon RX 7900 XT/7900 XTX/7900 GRE/7900MOllamaQ4_K_M24.6 tok/s24.6 tok/s20.5 GB65,536 tokens85.31
Apple M5 MaxLM Studio—24.3 tok/s20.8 tok/sn/a123,904 tokens73.03
Apple M5 ProLM Studio—17.1 tok/s15.2 tok/sn/a65,536 tokens66.72
AMD Radeon RX 9070 XTOllamaQ4_K_M15.0 tok/s15.0 tok/s16.9 GB8,192 tokens85.01
Apple M4 ProLM Studio—14.0 tok/s14.0 tok/sn/a32,768 tokens85.71
Apple M5LM Studio—7.6 tok/s7.6 tok/sn/a8,192 tokens43.11

Benchmark runs

All 26 qwen3.8-27b runs submitted so far — expand one for its hardware, quality breakdown and per-scenario detail.

Frequently asked questions

Is qwen3.8-27b good for coding?
In our benchmarks, qwen3.8-27b scores 51.4/100 for coding. It runs at about 50.3 tok/s, so if you want more speed, Tiel-Coder-35B-A3B-MLX-oQ4e is faster (~107.9 tok/s) and still scores well for coding (81.5/100).
Is qwen3.8-27b good for agentic (tool-using) tasks?
In our benchmarks, qwen3.8-27b scores 84.7/100 for agentic workflows. It runs at about 50.3 tok/s, so if you want more speed, Nex-N2.5-mini-APEX-Mini is faster (~88.2 tok/s) and still scores well for agentic workflows (91.3/100).
How fast is qwen3.8-27b for local inference?
Across 26 community benchmark runs, qwen3.8-27b reaches up to 156.8 tok/s and averages 50.3 tok/s, with the fastest results on NVIDIA GeForce RTX 5090.
How much memory does qwen3.8-27b need?
The leanest observed configuration used about 16.3 GB of memory (quantizations tested: Q4_K_M).
Which tools have been used to run qwen3.8-27b?
Benchmarks were submitted using LM Studio, Ollama, llama.cpp, ninfer, vLLM. Results are community-contributed and updated as new runs arrive.