Fastest & Best quality & Largest context
gpt-oss:20b
20.9B
64.0 tok/sQuality 74.8Context 65,536 tokens
Highest token generation speed at 64.0 tok/s. Highest completed quality score at 74.8. Largest context that fits in 15GB VRAM, at 65,536 tokens.
Hardware Benchmarks
Compare RTX 5080 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for RTX 5080.
Fastest & Best quality & Largest context
20.9B
Highest token generation speed at 64.0 tok/s. Highest completed quality score at 74.8. Largest context that fits in 15GB VRAM, at 65,536 tokens.
Which models actually run best on the RTX 5080, by task, from community benchmark data.
Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5080.