Hardware Benchmarks

Best Local LLM Performance for 8× RTX 4090 D

Compare 8× RTX 4090 D benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

8× NVIDIA GeForce RTX 4090 D
VRAM
384GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for 8× RTX 4090 D.

Fastest & Best quality

qwen3.8-flash-next

Unknown

174.8 tok/sQuality 79.6

Highest token generation speed at 174.8 tok/s. Highest completed quality score at 79.6.

Frequently asked questions

Which models actually run best on the 8× RTX 4090 D, by task, from community benchmark data.

What is the best local model for coding on the 8× RTX 4090 D?
On the 8× RTX 4090 D, qwen3.8-flash-next ranks first for coding in our benchmarks (51.1/100 at ~173.4 tok/s).
What is the best local model for agentic workflows on the 8× RTX 4090 D?
On the 8× RTX 4090 D, qwen3.8-flash-next ranks first for agentic workflows in our benchmarks (81.5/100 at ~173.4 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for 8× NVIDIA GeForce RTX 4090 D.