Hardware Benchmarks

Best Local LLM Performance for 2× L40s

Compare 2× L40s benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

2× NVIDIA L40S
VRAM
88GB
Architecture
Multi-GPU

Recommended Models

Models we recommend for 2× L40s.

Fastest & Best quality & Largest context

gpt-oss:20b

20.9B

120.7 tok/sQuality 76.5Context 8,192 tokens

Highest token generation speed at 120.7 tok/s. Highest completed quality score at 76.5. Largest context that fits in 88GB VRAM, at 8,192 tokens.

Frequently asked questions

Which models actually run best on the 2× L40s, by task, from community benchmark data.

What is the best local model for coding on the 2× L40s?
On the 2× L40s, gpt-oss:20b ranks first for coding in our benchmarks (65.0/100 at ~120.7 tok/s).
What is the best local model for agentic workflows on the 2× L40s?
On the 2× L40s, gpt-oss:20b ranks first for agentic workflows in our benchmarks (85.3/100 at ~120.7 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for 2× NVIDIA L40S.