Fastest & Largest context
unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q8_K_XL
26B
125.2 tok/sContext 131,072 tokens
Highest token generation speed at 125.2 tok/s. Largest context that fits in 20GB VRAM, at 131,072 tokens.
Hardware Benchmarks
Compare RTX 3080 benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for RTX 3080.
Fastest & Largest context
26B
Highest token generation speed at 125.2 tok/s. Largest context that fits in 20GB VRAM, at 131,072 tokens.
Best quality
27B
Highest completed quality score at 90.1.
Completed quality-ranked benchmark results for NVIDIA GeForce RTX 3080.