Hardware Benchmarks

Best Local LLM Performance for 2× RTX 4090

Compare 2× RTX 4090 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

2× NVIDIA GeForce RTX 4090
VRAM
48GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for 2× RTX 4090.

Fastest & Best quality & Largest context

jerry0012000/DeepSeek-R1-Distill-Qwen-32B-Q4_K_M:latest

32.8B

42.6 tok/sQuality 57.1Context 8,192 tokens

Highest token generation speed at 42.6 tok/s. Highest completed quality score at 57.1. Largest context that fits in 48GB VRAM, at 8,192 tokens.

Frequently asked questions

Which models actually run best on the 2× RTX 4090, by task, from community benchmark data.

What is the best local model for coding on the 2× RTX 4090?
On the 2× RTX 4090, jerry0012000/DeepSeek-R1-Distill-Qwen-32B-Q4_K_M:latest ranks first for coding in our benchmarks (51.2/100 at ~42.6 tok/s).
What is the best local model for agentic workflows on the 2× RTX 4090?
On the 2× RTX 4090, jerry0012000/DeepSeek-R1-Distill-Qwen-32B-Q4_K_M:latest ranks first for agentic workflows in our benchmarks (74.1/100 at ~42.6 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for 2× NVIDIA GeForce RTX 4090.