Model benchmarks
Qwen3.6-35B-A3B-UD-IQ2_XXS local LLM performance
As of October 2026, Qwen3.6-35B-A3B-UD-IQ2_XXS runs at up to 226.0 tok/s for local inference (best of 12 community benchmark runs across 1 GPU).
Model size
35B
Peak speed
226.0 tok/s
Average speed
209.8 tok/s
Avg PP
1486.7 tok/s
Min memory
14.4 GB
Max context
65,536 tokens
Avg output / run
17,477 tokens
Avg runtime / run
1m 27s
Avg quality
77.6
Benchmark runs
12
GPUs tested
1
Quality by task
Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 12 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.
| Task | P5 (low) | Avg | P95 (high) |
|---|---|---|---|
| Overall | 74.4 | 77.6 | 81.1 |
| Agent Workflow | 65.8 | 75.6 | 90.2 |
| Code Generation | 65.0 | 71.9 | 76.4 |
| Role Play & Narrative | 68.0 | 83.3 | 92.2 |
| Research & Analysis | 72.2 | 79.7 | 84.1 |
Performance by hardware and tool
Every hardware/tool/quantization combination Qwen3.6-35B-A3B-UD-IQ2_XXS has been benchmarked on, ranked by peak token generation speed. Last updated October 2026.
| Hardware | Tool | Quant | Best tok/s | Avg tok/s | Memory | Context | Quality | Runs |
|---|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | lm-studio | — | 226.0 tok/s | 209.8 tok/s | 14.4 GB | 65,536 tokens | 77.6 | 12 |
Benchmark runs
All 12 Qwen3.6-35B-A3B-UD-IQ2_XXS runs submitted so far — expand one for its hardware, quality breakdown and per-scenario detail.
Frequently asked questions
- How fast is Qwen3.6-35B-A3B-UD-IQ2_XXS for local inference?
- Across 12 community benchmark runs, Qwen3.6-35B-A3B-UD-IQ2_XXS reaches up to 226.0 tok/s and averages 209.8 tok/s, with the fastest results on AMD Radeon RX 7900 XTX.
- How much memory does Qwen3.6-35B-A3B-UD-IQ2_XXS need?
- The leanest observed configuration used about 14.4 GB of memory.
- Which tools have been used to run Qwen3.6-35B-A3B-UD-IQ2_XXS?
- Benchmarks were submitted using lm-studio. Results are community-contributed and updated as new runs arrive.