Hardware Benchmarks

Best Local LLM Performance for 2× RTX 3080

Compare 2× RTX 3080 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

2× NVIDIA GeForce RTX 3080
VRAM
20GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for 2× RTX 3080.

Fastest & Largest context

DavidAU/LFM2.5-2.6B-Qwen3.8-Turbo-Brilliance-Power-X12-NEO-MAX-GGUF

2.6B

176.0 tok/sContext 8,192 tokens

Highest token generation speed at 176.0 tok/s. Largest context that fits in 20GB VRAM, at 8,192 tokens.

Best quality

unsloth/gemma-4-26B-A4B-it-qat-GGUF

26B

40.8 tok/sQuality 70.9

Highest completed quality score at 70.9.

Frequently asked questions

Which models actually run best on the 2× RTX 3080, by task, from community benchmark data.

What is the best local model for coding on the 2× RTX 3080?
On the 2× RTX 3080, unsloth/gemma-4-26B-A4B-it-qat-GGUF ranks first for coding in our benchmarks (69.4/100 at ~40.8 tok/s).
What is the best local model for agentic workflows on the 2× RTX 3080?
On the 2× RTX 3080, unsloth/gemma-4-26B-A4B-it-qat-GGUF ranks first for agentic workflows in our benchmarks (80.4/100 at ~40.8 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for 2× NVIDIA GeForce RTX 3080.