Hardware Benchmarks

Best Local LLM Performance for RTX 5090

Compare RTX 5090 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 5090
VRAM
32GB

Recommended Models

Models we recommend for RTX 5090.

Fastest

qwen3-coder:30b

30.5B

240.2 tok/s

Highest token generation speed at 240.2 tok/s.

Best quality & Largest context

esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF

27B

134.1 tok/sQuality 86.9Context 65,536 tokens

Highest completed quality score at 86.9. Largest context that fits in 32GB VRAM, at 65,536 tokens.

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5090.