Hardware Benchmarks

Best Local LLM Performance for RTX 2050

Compare RTX 2050 benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA GeForce RTX 2050
VRAM
4GB

Recommended Models

Models we recommend for RTX 2050.

Fastest & Largest context

deepseek-r1:1.5b

1.8B

54.5 tok/sContext 8.192 tokens

Highest token generation speed at 54.5 tok/s. Largest context that fits in 4GB VRAM, at 8.192 tokens.

Best quality

qwen3-4b-abliterated-32k:latest

4.0B

17.2 tok/sQuality 61.4

Highest completed quality score at 61.4.

Frequently asked questions

Which models actually run best on the RTX 2050, by task, from community benchmark data.

What is the best local model for coding on the RTX 2050?
On the RTX 2050, qwen3-4b-abliterated-32k:latest ranks first for coding in our benchmarks (44.2/100 at ~17.2 tok/s).
What is the best local model for agentic workflows on the RTX 2050?
On the RTX 2050, qwen3-4b-abliterated-32k:latest ranks first for agentic workflows in our benchmarks (70.2/100 at ~17.2 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA GeForce RTX 2050.