Model benchmarks

openai/gpt-oss-20b:2 local LLM performance

As of August 2026, openai/gpt-oss-20b:2 runs at up to 158.7 tok/s for local inference (best of 2 community benchmark runs across 1 GPU).

LM Studio
ShareRedditX

Model size

20B

Peak speed

158.7 tok/s

Average speed

154.8 tok/s

Min memory

11.3 GB

Max context

2,915 tokens

Best quality

91.5

Benchmark runs

2

GPUs tested

1

Performance by hardware and tool

Every hardware/tool/quantization combination openai/gpt-oss-20b:2 has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
NVIDIA GeForce RTX 5070LM Studio158.7 tok/s154.8 tok/s11.3 GB2,915 tokens71.12

Frequently asked questions

How fast is openai/gpt-oss-20b:2 for local inference?
Across 2 community benchmark runs, openai/gpt-oss-20b:2 reaches up to 158.7 tok/s and averages 154.8 tok/s, with the fastest results on NVIDIA GeForce RTX 5070.
How much memory does openai/gpt-oss-20b:2 need?
The leanest observed configuration used about 11.3 GB of memory.
Which tools have been used to run openai/gpt-oss-20b:2?
Benchmarks were submitted using LM Studio. Results are community-contributed and updated as new runs arrive.