Hardware Benchmarks

Best Local LLM Performance for RTX 4090

Compare RTX 4090 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 4090
VRAM
23GB

Recommended Models

Models we recommend for RTX 4090.

Fastest & Best quality & Largest context

unsloth/Qwen3.8-27B-GGUF:IQ3_S

27B

107.2 tok/sQuality 88.4Context 131,072 tokens

Highest token generation speed at 107.2 tok/s. Highest completed quality score at 88.4. Largest context that fits in 23GB VRAM, at 131,072 tokens.

Frequently asked questions

Which models actually run best on the RTX 4090, by task, from community benchmark data.

What is the best local model for coding on the RTX 4090?
On the RTX 4090, unsloth/Qwen3.8-27B-GGUF:Q4_K_M ranks first for coding in our benchmarks (83.7/100 at ~89.2 tok/s). unsloth/Qwen3.8-27B-GGUF:IQ3_S is faster (~106.4 tok/s) and still scores well (80.5/100) — a good pick if you'd rather trade a little quality for speed.
What is the best local model for agentic workflows on the RTX 4090?
On the RTX 4090, unsloth/Qwen3.8-27B-GGUF:IQ3_S ranks first for agentic workflows in our benchmarks (90.4/100 at ~106.4 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 4090.