Model benchmarks
qwen3.8-27b local LLM performance
As of September 2026, qwen3.8-27b runs at up to 156.8 tok/s for local inference (best of 26 community benchmark runs across 12 GPUs).
Model size
27B
Peak speed
156.8 tok/s
Average speed
50.3 tok/s
Avg PP
1365.2 tok/s
Min memory
16.3 GB
Max context
123,904 tokens
Avg output / run
49,550 tokens
Avg runtime / run
24m 14s
Avg quality
73.1
Benchmark runs
26
GPUs tested
12
Quality by task
Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 26 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.
| Task | P5 (low) | Avg | P95 (high) |
|---|---|---|---|
| Overall | 44.7 | 73.1 | 88.0 |
| Agent Workflow | 74.3 | 84.7 | 91.1 |
| Code Generation | 1.6 | 51.4 | 82.9 |
| Role Play & Narrative | 79.8 | 89.0 | 94.5 |
| Research & Analysis | 0.3 | 67.4 | 90.3 |
Performance by hardware and tool
Every hardware/tool/quantization combination qwen3.8-27b has been benchmarked on, ranked by peak token generation speed. Last updated September 2026.
| Hardware | Tool | Quant | Best tok/s | Avg tok/s | Memory | Context | Quality | Runs |
|---|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | ninfer | — | 156.8 tok/s | 150.0 tok/s | n/a | 65,536 tokens | 68.0 | 2 |
| NVIDIA GeForce RTX 5090 | vLLM | — | 155.2 tok/s | 155.2 tok/s | n/a | 65,536 tokens | 89.1 | 1 |
| NVIDIA L4 | Ollama | Q4_K_M | 109.3 tok/s | 109.3 tok/s | 16.7 GB | 65,536 tokens | 86.7 | 1 |
| AMD Radeon RX 7900 XT/7900 XTX/7900M | llama.cpp | — | 52.4 tok/s | 51.8 tok/s | n/a | 65,536 tokens | 58.9 | 4 |
| AMD Radeon RX 7900 XT/7900 XTX/7900M | LM Studio | — | 51.9 tok/s | 51.9 tok/s | n/a | 65,536 tokens | 86.3 | 1 |
| AMD Radeon RX 7900 XTX | LM Studio | — | 51.8 tok/s | 51.8 tok/s | n/a | 65,536 tokens | 45.1 | 1 |
| NVIDIA GeForce RTX 5070 Ti | LM Studio | — | 46.9 tok/s | 42.8 tok/s | n/a | 121,072 tokens | 87.9 | 2 |
| AMD Radeon RX 7900 XTX | Ollama | Q4_K_M | 39.4 tok/s | 38.4 tok/s | 16.3 GB | 65,536 tokens | 82.2 | 4 |
| NVIDIA GeForce RTX 4090 | LM Studio | — | 39.0 tok/s | 39.0 tok/s | n/a | 8,192 tokens | 66.7 | 1 |
| AMD Radeon RX 7900 XT/7900 XTX/7900 GRE/7900M | Ollama | Q4_K_M | 24.6 tok/s | 24.6 tok/s | 20.5 GB | 65,536 tokens | 85.3 | 1 |
| Apple M5 Max | LM Studio | — | 24.3 tok/s | 20.8 tok/s | n/a | 123,904 tokens | 73.0 | 3 |
| Apple M5 Pro | LM Studio | — | 17.1 tok/s | 15.2 tok/s | n/a | 65,536 tokens | 66.7 | 2 |
| AMD Radeon RX 9070 XT | Ollama | Q4_K_M | 15.0 tok/s | 15.0 tok/s | 16.9 GB | 8,192 tokens | 85.0 | 1 |
| Apple M4 Pro | LM Studio | — | 14.0 tok/s | 14.0 tok/s | n/a | 32,768 tokens | 85.7 | 1 |
| Apple M5 | LM Studio | — | 7.6 tok/s | 7.6 tok/s | n/a | 8,192 tokens | 43.1 | 1 |
Benchmark runs
All 26 qwen3.8-27b runs submitted so far — expand one for its hardware, quality breakdown and per-scenario detail.
Frequently asked questions
- Is qwen3.8-27b good for coding?
- In our benchmarks, qwen3.8-27b scores 51.4/100 for coding. It runs at about 50.3 tok/s, so if you want more speed, Tiel-Coder-35B-A3B-MLX-oQ4e is faster (~107.9 tok/s) and still scores well for coding (81.5/100).
- Is qwen3.8-27b good for agentic (tool-using) tasks?
- In our benchmarks, qwen3.8-27b scores 84.7/100 for agentic workflows. It runs at about 50.3 tok/s, so if you want more speed, Nex-N2.5-mini-APEX-Mini is faster (~88.2 tok/s) and still scores well for agentic workflows (91.3/100).
- How fast is qwen3.8-27b for local inference?
- Across 26 community benchmark runs, qwen3.8-27b reaches up to 156.8 tok/s and averages 50.3 tok/s, with the fastest results on NVIDIA GeForce RTX 5090.
- How much memory does qwen3.8-27b need?
- The leanest observed configuration used about 16.3 GB of memory (quantizations tested: Q4_K_M).
- Which tools have been used to run qwen3.8-27b?
- Benchmarks were submitted using LM Studio, Ollama, llama.cpp, ninfer, vLLM. Results are community-contributed and updated as new runs arrive.