Hardware Benchmarks

Best Local LLM Performance for RTX 4080 Super

Compare RTX 4080 Super benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 4080 SUPER
VRAM
16GB

Recommended Models

Models we recommend for RTX 4080 Super.

Fastest

Qwen3.8-27B-UD-IQ3_XXS

27B

81.6 tok/s

Highest token generation speed at 81.6 tok/s.

Best quality & Largest context

Qwen3.8-27B-UD-IQ3_S

27B

77.3 tok/sQuality 86.4Context 100,000 tokens

Highest completed quality score at 86.4. Largest context that fits in 16GB VRAM, at 100,000 tokens.

Frequently asked questions

Which models actually run best on the RTX 4080 Super, by task, from community benchmark data.

What is the best local model for coding on the RTX 4080 Super?
On the RTX 4080 Super, Qwen3.8-27B-UD-IQ3_S ranks first for coding in our benchmarks (74.0/100 at ~78.1 tok/s).
What is the best local model for agentic workflows on the RTX 4080 Super?
On the RTX 4080 Super, Qwen3.8-27B-UD-IQ3_S ranks first for agentic workflows in our benchmarks (83.3/100 at ~78.1 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 4080 SUPER.