Fastest
qwen3.6-35b-a3b
35B
437.3 tok/s
Highest token generation speed at 437.3 tok/s.
Hardware Benchmarks
Compare RTX 5090 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for RTX 5090.
Fastest
35B
Highest token generation speed at 437.3 tok/s.
Best quality
27B
Highest completed quality score at 89.1.
Largest context
Unknown
Largest context that fits in 32GB VRAM, at 262,144 tokens.
Which models actually run best on the RTX 5090, by task, from community benchmark data.
Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5090.