Model benchmarks

qwen/qwen3-coder-30b local LLM performance

As of August 2026, qwen/qwen3-coder-30b runs at up to 240.2 tok/s for local inference (best of 2 community benchmark runs across 2 GPUs).

LM StudioOllamaQ4_K_M
ShareRedditX

Model size

30B

Peak speed

240.2 tok/s

Average speed

162.4 tok/s

Min memory

16.0 GB

Max context

32,768 tokens

Avg output / run

7,437 tokens

Avg runtime / run

1m 36s

Avg quality

55.4

Benchmark runs

2

GPUs tested

2

Quality by task

Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 2 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.

TaskP5 (low)AvgP95 (high)
Overall54.555.456.3
Agent Workflow41.643.846.1
Code Generation56.958.760.4
Role Play & Narrative61.261.762.1
Research & Analysis57.657.657.7

Performance by hardware and tool

Every hardware/tool/quantization combination qwen/qwen3-coder-30b has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
NVIDIA GeForce RTX 5090OllamaQ4_K_M240.2 tok/s240.2 tok/s18.9 GB32,768 tokens56.41
Apple M5 ProLM Studio84.5 tok/s84.5 tok/s16.0 GB32,768 tokens54.51

Frequently asked questions

Is qwen3-coder:30b good for coding?
In our benchmarks, qwen3-coder:30b scores 58.7/100 for coding. It runs at about 162.4 tok/s, so if you want more speed, Gemma4:E2B/QAT-MTP@131K is faster (~303.9 tok/s) and still scores well for coding (66.9/100).
Is qwen3-coder:30b good for agentic (tool-using) tasks?
In our benchmarks, qwen3-coder:30b scores 43.8/100 for agentic workflows. It runs at about 162.4 tok/s, so if you want more speed, Gemma4:E2B/QAT-MTP@131K is faster (~303.9 tok/s) and still scores well for agentic workflows (71.7/100).
How fast is qwen/qwen3-coder-30b for local inference?
Across 2 community benchmark runs, qwen/qwen3-coder-30b reaches up to 240.2 tok/s and averages 162.4 tok/s, with the fastest results on NVIDIA GeForce RTX 5090.
How much memory does qwen/qwen3-coder-30b need?
The leanest observed configuration used about 16.0 GB of memory (quantizations tested: Q4_K_M).
Which tools have been used to run qwen/qwen3-coder-30b?
Benchmarks were submitted using LM Studio, Ollama. Results are community-contributed and updated as new runs arrive.