Model benchmarks

qwen35_122b local LLM performance

As of August 2026, qwen35_122b runs at up to 15.2 tok/s for local inference (best of 5 community benchmark runs across 1 GPU).

LM Studio
ShareRedditX

Model size

122B

Peak speed

15.2 tok/s

Average speed

15.2 tok/s

Min memory

59.6 GB

Max context

6,674 tokens

Best quality

88.4

Benchmark runs

5

GPUs tested

1

Performance by hardware and tool

Every hardware/tool/quantization combination qwen35_122b has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
NVIDIA GeForce RTX 3080LM Studio15.2 tok/s15.2 tok/s59.6 GB6,674 tokens66.75

Frequently asked questions

How fast is qwen35_122b for local inference?
Across 5 community benchmark runs, qwen35_122b reaches up to 15.2 tok/s and averages 15.2 tok/s, with the fastest results on NVIDIA GeForce RTX 3080.
How much memory does qwen35_122b need?
The leanest observed configuration used about 59.6 GB of memory.
Which tools have been used to run qwen35_122b?
Benchmarks were submitted using LM Studio. Results are community-contributed and updated as new runs arrive.