Hardware Benchmarks

Best Local LLM Performance for RTX 5090 Laptop

Compare RTX 5090 Laptop benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 5090 Laptop GPU
VRAM
23GB

Recommended Models

Models we recommend for RTX 5090 Laptop.

Fastest & Best quality

Qwen3.8-27B-Q8-DFlash2

27B

69.9 tok/sQuality 86.6

Highest token generation speed at 69.9 tok/s. Highest completed quality score at 86.6.

Largest context

Qwen3.8-27B-Q4-MTP

27B

48.2 tok/sContext 8,192 tokens

Largest context that fits in 23GB VRAM, at 8,192 tokens.

Frequently asked questions

Which models actually run best on the RTX 5090 Laptop, by task, from community benchmark data.

What is the best local model for coding on the RTX 5090 Laptop?
On the RTX 5090 Laptop, Qwen3.8-27B-Q4-MTP ranks first for coding in our benchmarks (83.5/100 at ~48.2 tok/s). Qwen3.8-27B-Q8-DFlash2 is faster (~69.9 tok/s) and still scores well (82.4/100) — a good pick if you'd rather trade a little quality for speed.
What is the best local model for agentic workflows on the RTX 5090 Laptop?
On the RTX 5090 Laptop, Qwen3.8-27B-Q4-MTP ranks first for agentic workflows in our benchmarks (87.7/100 at ~48.2 tok/s). Qwen3.8-27B-Q8-DFlash2 is faster (~69.9 tok/s) and still scores well (84.2/100) — a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5090 Laptop GPU.