Fastest benchmarked model
gpt-oss:20b
20.9B
120.7 tok/s
This benchmark delivered the highest token generation speed at 120.7 tok/s.
Hardware Benchmarks
Compare L40s benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for L40s.
Fastest benchmarked model
20.9B
This benchmark delivered the highest token generation speed at 120.7 tok/s.
Best-quality benchmarked model
32B
This benchmark has the highest completed quality score at 58.7.
Completed quality-ranked benchmark results for NVIDIA L40S.