Hardware Benchmarks

Best Local LLM Performance for RTX 5080

Compare RTX 5080 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 5080
VRAM
15GB

Recommended Models

Models we recommend for RTX 5080.

Fastest & Best quality & Largest context

gpt-oss:20b

20.9B

64.0 tok/sQuality 74.8Context 65,536 tokens

Highest token generation speed at 64.0 tok/s. Highest completed quality score at 74.8. Largest context that fits in 15GB VRAM, at 65,536 tokens.

Frequently asked questions

Which models actually run best on the RTX 5080, by task, from community benchmark data.

What is the best local model for coding on the RTX 5080?
On the RTX 5080, gpt-oss:20b ranks first for coding in our benchmarks (54.6/100 at ~35.0 tok/s).
What is the best local model for agentic workflows on the RTX 5080?
On the RTX 5080, qwen3.5:9b ranks first for agentic workflows in our benchmarks (70.8/100 at ~22.3 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5080.