Model benchmarks
qwen/qwen3.6-35b-a3b local LLM performance
As of August 2026, qwen/qwen3.6-35b-a3b runs at up to 437.3 tok/s for local inference (best of 12 community benchmark runs across 5 GPUs).
Model size
35B
Peak speed
437.3 tok/s
Average speed
70.0 tok/s
Min memory
14.4 GB
Max context
131.072 tokens
Avg output / run
23.114 tokens
Avg runtime / run
6m 1s
Avg quality
73.1
Benchmark runs
12
GPUs tested
5
Quality by task
Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 12 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.
| Task | P5 (low) | Avg | P95 (high) |
|---|---|---|---|
| Overall | 56.9 | 73.1 | 83.6 |
| Agent Workflow | 35.8 | 70.1 | 85.9 |
| Code Generation | 28.0 | 60.6 | 79.5 |
| Role Play & Narrative | 72.1 | 82.7 | 92.2 |
| Research & Analysis | 68.0 | 78.8 | 88.6 |
Performance by hardware and tool
Every hardware/tool/quantization combination qwen/qwen3.6-35b-a3b has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.
| Hardware | Tool | Quant | Best tok/s | Avg tok/s | Memory | Context | Quality | Runs |
|---|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5090 | vLLM | — | 437.3 tok/s | 437.3 tok/s | 26.5 GB | 65.536 tokens | 84.7 | 1 |
| NVIDIA GeForce RTX 5070 Ti | LM Studio | — | 75.9 tok/s | 75.9 tok/s | 14.4 GB | 121.072 tokens | 70.8 | 1 |
| Apple M4 Max | LM Studio | — | 69.8 tok/s | 69.8 tok/s | 35.2 GB | 120.000 tokens | 72.9 | 1 |
| AMD Radeon RX 7900 XTX | LM Studio | — | 46.1 tok/s | 31.1 tok/s | 17.1 GB | 65.536 tokens | 68.1 | 6 |
| CPU only | LM Studio | — | 25.0 tok/s | 24.9 tok/s | 17.1 GB | 131.072 tokens | 78.7 | 2 |
| AMD Radeon RX 7600 XT | LM Studio | — | 20.8 tok/s | 20.8 tok/s | 21.1 GB | 32.768 tokens | 82.2 | 1 |
Benchmark runs
All 12 qwen/qwen3.6-35b-a3b runs submitted so far — expand one for its hardware, quality breakdown and per-scenario detail.
Frequently asked questions
- Is qwen/qwen3.6-35b-a3b good for coding?
- In our benchmarks, qwen/qwen3.6-35b-a3b scores 60.6/100 for coding. It runs at about 70.0 tok/s, so if you want more speed, unsloth/Qwen3.8-27B-GGUF:IQ3_S is faster (~104.3 tok/s) and still scores well for coding (83.0/100).
- Is qwen/qwen3.6-35b-a3b good for agentic (tool-using) tasks?
- In our benchmarks, qwen/qwen3.6-35b-a3b scores 70.1/100 for agentic workflows. It runs at about 70.0 tok/s, so if you want more speed, Tiel-Coder-35B-A3B-Q4_K_S is faster (~165.0 tok/s) and still scores well for agentic workflows (89.3/100).
- How fast is qwen/qwen3.6-35b-a3b for local inference?
- Across 12 community benchmark runs, qwen/qwen3.6-35b-a3b reaches up to 437.3 tok/s and averages 70.0 tok/s, with the fastest results on NVIDIA GeForce RTX 5090.
- How much memory does qwen/qwen3.6-35b-a3b need?
- The leanest observed configuration used about 14.4 GB of memory.
- Which tools have been used to run qwen/qwen3.6-35b-a3b?
- Benchmarks were submitted using LM Studio, vLLM. Results are community-contributed and updated as new runs arrive.