Hardware Benchmarks

Best Local LLM Performance for L40s

Compare L40s benchmark results and see which models actually earn the best completed quality scores.

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA L40S
VRAM
44GB

Recommended Models

Models we recommend for L40s.

Fastest benchmarked model

gpt-oss:20b

20.9B

120.7 tok/s

This benchmark delivered the highest token generation speed at 120.7 tok/s.

Best-quality benchmarked model

vllm-32b

32B

37.5 tok/sQuality 58.7

This benchmark has the highest completed quality score at 58.7.

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA L40S.