Hardware Benchmarks

Best Local LLM Performance for RTX 5060 Ti + RTX 5080

Compare RTX 5060 Ti + RTX 5080 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 5060 Ti + NVIDIA GeForce RTX 5080
VRAM
32GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for RTX 5060 Ti + RTX 5080.

Fastest

gemma3:270m

268.10M

568.1 tok/s

Highest token generation speed at 568.1 tok/s.

Best quality

bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L

7B

40.2 tok/sQuality 80.8

Highest completed quality score at 80.8.

Largest context

bench-qwen3.6-27b:latest

27.3B

25.1 tok/sContext 65,536 tokens

Largest context that fits in 32GB VRAM, at 65,536 tokens.

Frequently asked questions

Which models actually run best on the RTX 5060 Ti + RTX 5080, by task, from community benchmark data.

What is the best local model for coding on the RTX 5060 Ti + RTX 5080?
On the RTX 5060 Ti + RTX 5080, bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L ranks first for coding in our benchmarks (74.0/100 at ~38.3 tok/s).
What is the best local model for agentic workflows on the RTX 5060 Ti + RTX 5080?
On the RTX 5060 Ti + RTX 5080, bench-qwen3.6-27b:latest ranks first for agentic workflows in our benchmarks (89.4/100 at ~25.1 tok/s). qwen3.6:latest is faster (~142.1 tok/s) and still scores well (87.0/100) — a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 5060 Ti + NVIDIA GeForce RTX 5080.