Model benchmarks
bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L local LLM performance
As of August 2026, bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L runs at up to 40.2 tok/s for local inference (best of 2 community benchmark runs across 1 GPU).
llama.cppQ6_K
Model size
7B
Peak speed
40.2 tok/s
Average speed
38.3 tok/s
Min memory
3.4 GB
Max context
4,687 tokens
Best quality
89.6
Benchmark runs
2
GPUs tested
1
Performance by hardware and tool
Every hardware/tool/quantization combination bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.
| Hardware | Tool | Quant | Best tok/s | Avg tok/s | Memory | Context | Quality | Runs |
|---|---|---|---|---|---|---|---|---|
| NVIDIA GeForce RTX 5080 | llama.cpp | Q6_K | 40.2 tok/s | 38.3 tok/s | 3.4 GB | 4,687 tokens | 79.8 | 2 |
Frequently asked questions
- How fast is bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L for local inference?
- Across 2 community benchmark runs, bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L reaches up to 40.2 tok/s and averages 38.3 tok/s, with the fastest results on NVIDIA GeForce RTX 5080.
- How much memory does bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L need?
- The leanest observed configuration used about 3.4 GB of memory (quantizations tested: Q6_K).
- Which tools have been used to run bartowski/Qwen_Qwen3-Coder-Next-GGUF:Q6_K_L?
- Benchmarks were submitted using llama.cpp. Results are community-contributed and updated as new runs arrive.