Hardware Benchmarks

Best Local LLM Performance for 3× RTX 5090

Compare 3× RTX 5090 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

3× NVIDIA GeForce RTX 5090
VRAM
93GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for 3× RTX 5090.

Fastest

Sphinsikus-Chronist-31B-Q4_K_M

31B

56.7 tok/s

Highest token generation speed at 56.7 tok/s.

Best quality & Largest context

gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4

31B

55.2 tok/sQuality 76.8Context 8,192 tokens

Highest completed quality score at 76.8. Largest context that fits in 93GB VRAM, at 8,192 tokens.

Frequently asked questions

Which models actually run best on the 3× RTX 5090, by task, from community benchmark data.

What is the best local model for coding on the 3× RTX 5090?
On the 3× RTX 5090, gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4 ranks first for coding in our benchmarks (66.8/100 at ~55.2 tok/s).
What is the best local model for agentic workflows on the 3× RTX 5090?
On the 3× RTX 5090, gemma-4-31B-it-qat-q4_0-uncensored-heretic-NVFP4 ranks first for agentic workflows in our benchmarks (81.8/100 at ~55.2 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for 3× NVIDIA GeForce RTX 5090.