Fastest
openai/gpt-oss-20b:2
20B
158.7 tok/s
Highest token generation speed at 158.7 tok/s.
Hardware Benchmarks
Compare RTX 5060 + RTX 5070 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for RTX 5060 + RTX 5070.
Fastest
20B
Highest token generation speed at 158.7 tok/s.
Best quality
9B
Highest completed quality score at 74.0.
Largest context
20B
Largest context that fits in 18GB VRAM, at 65,536 tokens.
Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5060 + NVIDIA GeForce RTX 5070.