Hardware Benchmarks

Best Local LLM Performance for RTX 4070 Ti Super

Compare RTX 4070 Ti Super benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 4070 Ti SUPER
VRAM
15GB

Recommended Models

Models we recommend for RTX 4070 Ti Super.

Fastest

lmstudio-community/lfm2.5-2.6b

2.6B

165.9 tok/s

Highest token generation speed at 165.9 tok/s.

Best quality

ornith-1.0-35b

35B

21.3 tok/sQuality 79.3

Highest completed quality score at 79.3.

Largest context

nanbeige4.1-3b

3B

70.0 tok/sContext 65,536 tokens

Largest context that fits in 15GB VRAM, at 65,536 tokens.

Frequently asked questions

Which models actually run best on the RTX 4070 Ti Super, by task, from community benchmark data.

What is the best local model for coding on the RTX 4070 Ti Super?
On the RTX 4070 Ti Super, ornith-1.0-35b ranks first for coding in our benchmarks (82.0/100 at ~21.3 tok/s).
What is the best local model for agentic workflows on the RTX 4070 Ti Super?
On the RTX 4070 Ti Super, unsloth/gemma-4-26b-a4b-it-qat ranks first for agentic workflows in our benchmarks (86.3/100 at ~48.4 tok/s). lmstudio-community/lfm2.5-2.6b is faster (~132.7 tok/s) and still scores well (81.6/100) — a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 4070 Ti SUPER.