Hardware Benchmarks

Best Local LLM Performance for 2× RTX 4070

Compare 2× RTX 4070 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

2× NVIDIA GeForce RTX 4070
VRAM
24GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for 2× RTX 4070.

Fastest

ornith-1.5-9b-obliterated

9B

89.3 tok/s

Highest token generation speed at 89.3 tok/s.

Best quality & Largest context

Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP

27B

56.9 tok/sQuality 82.5Context 65,536 tokens

Highest completed quality score at 82.5. Largest context that fits in 24GB VRAM, at 65,536 tokens.

Frequently asked questions

Which models actually run best on the 2× RTX 4070, by task, from community benchmark data.

What is the best local model for coding on the 2× RTX 4070?
On the 2× RTX 4070, Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP ranks first for coding in our benchmarks (89.5/100 at ~56.9 tok/s).
What is the best local model for agentic workflows on the 2× RTX 4070?
On the 2× RTX 4070, ternary-bonsai-2-27b ranks first for agentic workflows in our benchmarks (90.0/100 at ~28.1 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for 2× NVIDIA GeForce RTX 4070.