Hardware Benchmarks

Best Local LLM Performance for L4 + RTX Pro 6000 Blackwell Server Edition

Compare L4 + RTX Pro 6000 Blackwell Server Edition benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA L4 + NVIDIA RTX PRO 6000 Blackwell Server Edition
VRAM
118GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for L4 + RTX Pro 6000 Blackwell Server Edition.

Fastest

nemotron-3.5-lightning:latest

32.9B

311.5 tok/s

Highest token generation speed at 311.5 tok/s.

Best quality & Largest context

qwen3.8:27b

27.3B

109.3 tok/sQuality 86.7Context 65,536 tokens

Highest completed quality score at 86.7. Largest context that fits in 118GB VRAM, at 65,536 tokens.

Frequently asked questions

Which models actually run best on the L4 + RTX Pro 6000 Blackwell Server Edition, by task, from community benchmark data.

What is the best local model for coding on the L4 + RTX Pro 6000 Blackwell Server Edition?
On the L4 + RTX Pro 6000 Blackwell Server Edition, qwen3.8:27b ranks first for coding in our benchmarks (76.4/100 at ~109.3 tok/s).
What is the best local model for agentic workflows on the L4 + RTX Pro 6000 Blackwell Server Edition?
On the L4 + RTX Pro 6000 Blackwell Server Edition, ornith-1.5:35b ranks first for agentic workflows in our benchmarks (90.4/100 at ~221.7 tok/s). nemotron-3.5-lightning:latest is faster (~311.5 tok/s) and still scores well (86.0/100) β€” a good pick if you'd rather trade a little quality for speed.

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA L4 + NVIDIA RTX PRO 6000 Blackwell Server Edition.