Fastest & Best quality & Largest context
gpt-oss:20b
20.9B
120.7 tok/sQuality 76.5Context 8,192 tokens
Highest token generation speed at 120.7 tok/s. Highest completed quality score at 76.5. Largest context that fits in 88GB VRAM, at 8,192 tokens.
Hardware Benchmarks
Compare 2× L40s benchmark results and see which models actually earn the best completed quality scores.
GPU-specific specifications for local LLM planning.
Models we recommend for 2× L40s.
Fastest & Best quality & Largest context
20.9B
Highest token generation speed at 120.7 tok/s. Highest completed quality score at 76.5. Largest context that fits in 88GB VRAM, at 8,192 tokens.
Which models actually run best on the 2× L40s, by task, from community benchmark data.
Completed quality-ranked benchmark results for 2× NVIDIA L40S.