Hardware Benchmarks

Best Local LLM Performance for RTX Pro 6000 Blackwell Max Q Workstation Edition

Compare RTX Pro 6000 Blackwell Max Q Workstation Edition benchmark results and see which models actually earn the best completed quality scores.

ShareRedditX

Hardware Overview

GPU-specific specifications for local LLM planning.

NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
VRAM
96GB

Recommended Models

Models we recommend for RTX Pro 6000 Blackwell Max Q Workstation Edition.

Fastest & Best quality & Largest context

/models/hub/models--unsloth--Qwen3.8-27B-GGUF/snapshots/f1bfb127c64f7072bdd2cad55f258b9c8b2910fe/Qwen3.8-27B-UD-Q6_K_XL.gguf

27B

90.1 tok/sQuality 88.6Context 250,000 tokens

Highest token generation speed at 90.1 tok/s. Highest completed quality score at 88.6. Largest context that fits in 96GB VRAM, at 250,000 tokens.

Frequently asked questions

Which models actually run best on the RTX Pro 6000 Blackwell Max Q Workstation Edition, by task, from community benchmark data.

What is the best local model for coding on the RTX Pro 6000 Blackwell Max Q Workstation Edition?
On the RTX Pro 6000 Blackwell Max Q Workstation Edition, /models/hub/models--unsloth--Qwen3.8-27B-GGUF/snapshots/f1bfb127c64f7072bdd2cad55f258b9c8b2910fe/Qwen3.8-27B-UD-Q6_K_XL.gguf ranks first for coding in our benchmarks (75.9/100 at ~84.3 tok/s).
What is the best local model for agentic workflows on the RTX Pro 6000 Blackwell Max Q Workstation Edition?
On the RTX Pro 6000 Blackwell Max Q Workstation Edition, /models/hub/models--unsloth--Qwen3.8-27B-GGUF/snapshots/f1bfb127c64f7072bdd2cad55f258b9c8b2910fe/Qwen3.8-27B-UD-Q6_K_XL.gguf ranks first for agentic workflows in our benchmarks (76.9/100 at ~84.3 tok/s).

Benchmark Results

Completed quality-ranked benchmark results for NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition.