Model benchmarks
gpt-oss:20b local LLM performance
As of August 2026, gpt-oss:20b runs at up to 155.2 tok/s for local inference (best of 30 community benchmark runs across 8 GPUs).
LM StudioOllamaMXFP4
Model size
20.9B
Peak speed
155.2 tok/s
Average speed
76.4 tok/s
Min memory
9.8 GB
Max context
256,000 tokens
Avg output / run
6,412 tokens
Avg runtime / run
9m 32s
Avg quality
69.5
Benchmark runs
30
GPUs tested
8
Quality by task
Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 30 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.
| Task | P5 (low) | Avg | P95 (high) |
|---|---|---|---|
| Overall | 55.7 | 69.5 | 76.2 |
| Agent Workflow | 37.6 | 76.8 | 90.4 |
| Code Generation | 39.6 | 61.5 | 77.2 |
| Role Play & Narrative | 55.1 | 66.9 | 76.3 |
| Research & Analysis | 58.8 | 72.7 | 81.8 |
Performance by hardware and tool
Every hardware/tool/quantization combination gpt-oss:20b has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.
| Hardware | Tool | Quant | Best tok/s | Avg tok/s | Memory | Context | Quality | Runs |
|---|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 4070 Ti SUPER | LM Studio | — | 155.2 tok/s | 151.7 tok/s | 11.3 GB | 16,384 tokens | 70.8 | 2 |
| NVIDIA L40S | Ollama | MXFP4 | 120.7 tok/s | 120.7 tok/s | 12.0 GB | 8,192 tokens | 76.5 | 1 |
| AMD Radeon RX 7900 XTX | LM Studio | — | 111.1 tok/s | 109.1 tok/s | 9.8 GB | 256,000 tokens | 67.0 | 9 |
| AMD Radeon RX 7900 XTX | Ollama | MXFP4 | 70.6 tok/s | 62.0 tok/s | 13.1 GB | 65,536 tokens | 73.5 | 11 |
| NVIDIA GeForce RTX 5080 | Ollama | MXFP4 | 64.0 tok/s | 35.0 tok/s | 13.3 GB | 65,536 tokens | 62.4 | 2 |
| Apple M5 | LM Studio | — | 45.3 tok/s | 45.3 tok/s | 9.8 GB | 8,192 tokens | 74.0 | 1 |
| NVIDIA GeForce RTX 5070 | LM Studio | — | 34.8 tok/s | 34.8 tok/s | 11.3 GB | 65,536 tokens | 72.8 | 1 |
| Apple M1 Pro | Ollama | MXFP4 | 31.1 tok/s | 31.1 tok/s | 11.9 GB | 8,192 tokens | 40.7 | 1 |
| NVIDIA GeForce RTX 5070 | Ollama | MXFP4 | 15.4 tok/s | 15.4 tok/s | 14.6 GB | 65,536 tokens | 68.2 | 1 |
| Apple M3 | Ollama | MXFP4 | 8.4 tok/s | 8.4 tok/s | 13.0 GB | 8,192 tokens | 74.3 | 1 |
Frequently asked questions
- Is gpt-oss:20b good for coding?
- In our benchmarks, gpt-oss:20b scores 61.5/100 for coding. It runs at about 76.4 tok/s, so if you want more speed, Gemma4:E2B/QAT-MTP@131K is faster (~303.9 tok/s) and still scores well for coding (66.9/100).
- Is gpt-oss:20b good for agentic (tool-using) tasks?
- In our benchmarks, gpt-oss:20b scores 76.8/100 for agentic workflows. It runs at about 76.4 tok/s, so if you want more speed, Gemma4:E2B/QAT-MTP@131K is faster (~303.9 tok/s) and still scores well for agentic workflows (71.7/100).
- How fast is gpt-oss:20b for local inference?
- Across 30 community benchmark runs, gpt-oss:20b reaches up to 155.2 tok/s and averages 76.4 tok/s, with the fastest results on NVIDIA GeForce RTX 4070 Ti SUPER.
- How much memory does gpt-oss:20b need?
- The leanest observed configuration used about 9.8 GB of memory (quantizations tested: MXFP4).
- Which tools have been used to run gpt-oss:20b?
- Benchmarks were submitted using LM Studio, Ollama. Results are community-contributed and updated as new runs arrive.