Fastest
qwen3-coder:30b
30.5B
240.2 tok/s
Highest token generation speed at 240.2 tok/s.
Hardware Benchmarks
Compare RTX 5090 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for RTX 5090.
Fastest
30.5B
Highest token generation speed at 240.2 tok/s.
Best quality & Largest context
27B
Highest completed quality score at 86.9. Largest context that fits in 32GB VRAM, at 65,536 tokens.
Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5090.