Hardware Benchmarks

Best Local LLM Performance for RTX 5090

Compare RTX 5090 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 5090
VRAM
32GB

Recommended Models

Models we recommend for RTX 5090.

Fastest

qwen3.6-35b-a3b

35B

437.3 tok/s

Highest token generation speed at 437.3 tok/s.

Best quality

qwen3.8-27b

27B

155.2 tok/sQuality 89.1

Highest completed quality score at 89.1.

Largest context

swift-1.5-iq3_xxs

Unknown

116.5 tok/sContext 262,144 tokens

Largest context that fits in 32GB VRAM, at 262,144 tokens.

Frequently asked questions

Which models actually run best on the RTX 5090, by task, from community benchmark data.

What is the best local model for coding on the RTX 5090?
On the RTX 5090, qwen3.6-35b-a3b ranks first for coding in our benchmarks (84.4/100 at ~437.3 tok/s).
What is the best local model for agentic workflows on the RTX 5090?
On the RTX 5090, esatapedico/Qwen3.8-27B-NVFP4-MTP-GGUF ranks first for agentic workflows in our benchmarks (93.0/100 at ~134.1 tok/s). mudler/Ornith-1.5-35B-A3B-APEX-MTP-GGUF is faster (~190.9 tok/s) and still scores well (85.7/100) β€” a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5090.