Model benchmarks

Qwen3.6-35B-A3B-UD-IQ2_XXS local LLM performance

As of October 2026, Qwen3.6-35B-A3B-UD-IQ2_XXS runs at up to 226.0 tok/s for local inference (best of 12 community benchmark runs across 1 GPU).

lm-studio
ShareRedditX

Model size

35B

Peak speed

226.0 tok/s

Average speed

209.8 tok/s

Avg PP

1486.7 tok/s

Min memory

14.4 GB

Max context

65,536 tokens

Avg output / run

17,477 tokens

Avg runtime / run

1m 27s

Avg quality

77.6

Benchmark runs

12

GPUs tested

1

Quality by task

Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 12 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.

TaskP5 (low)AvgP95 (high)
Overall74.477.681.1
Agent Workflow65.875.690.2
Code Generation65.071.976.4
Role Play & Narrative68.083.392.2
Research & Analysis72.279.784.1

Performance by hardware and tool

Every hardware/tool/quantization combination Qwen3.6-35B-A3B-UD-IQ2_XXS has been benchmarked on, ranked by peak token generation speed. Last updated October 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
AMD Radeon RX 7900 XTXlm-studio—226.0 tok/s209.8 tok/s14.4 GB65,536 tokens77.612

Benchmark runs

All 12 Qwen3.6-35B-A3B-UD-IQ2_XXS runs submitted so far — expand one for its hardware, quality breakdown and per-scenario detail.

Frequently asked questions

How fast is Qwen3.6-35B-A3B-UD-IQ2_XXS for local inference?
Across 12 community benchmark runs, Qwen3.6-35B-A3B-UD-IQ2_XXS reaches up to 226.0 tok/s and averages 209.8 tok/s, with the fastest results on AMD Radeon RX 7900 XTX.
How much memory does Qwen3.6-35B-A3B-UD-IQ2_XXS need?
The leanest observed configuration used about 14.4 GB of memory.
Which tools have been used to run Qwen3.6-35B-A3B-UD-IQ2_XXS?
Benchmarks were submitted using lm-studio. Results are community-contributed and updated as new runs arrive.