Model benchmarks
glm-4.7-flash:latest local LLM performance
As of April 2026, glm-4.7-flash:latest runs at up to 46.1 tok/s for local inference (best of 8 community benchmark runs across 1 GPU).
LM StudioOllamaQ4_K_M
Model size
29.9B
Peak speed
46.1 tok/s
Average speed
29.6 tok/s
Min memory
n/a
Max context
32,768 tokens
Best quality
90.9
Benchmark runs
8
GPUs tested
1
Performance by hardware and tool
Every hardware/tool/quantization combination glm-4.7-flash:latest has been benchmarked on, ranked by peak token generation speed. Last updated April 2026.
| Hardware | Tool | Quant | Best tok/s | Avg tok/s | Memory | Context | Quality | Runs |
|---|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | LM Studio | — | 46.1 tok/s | 45.6 tok/s | n/a | 32,768 tokens | 47.4 | 3 |
| AMD Radeon RX 7900 XTX | Ollama | Q4_K_M | 37.5 tok/s | 19.9 tok/s | 18.6 GB | 7,615 tokens | 66.9 | 5 |
Frequently asked questions
- How fast is glm-4.7-flash:latest for local inference?
- Across 8 community benchmark runs, glm-4.7-flash:latest reaches up to 46.1 tok/s and averages 29.6 tok/s, with the fastest results on AMD Radeon RX 7900 XTX.
- How much memory does glm-4.7-flash:latest need?
- The leanest observed configuration used about n/a of memory (quantizations tested: Q4_K_M).
- Which tools have been used to run glm-4.7-flash:latest?
- Benchmarks were submitted using LM Studio, Ollama. Results are community-contributed and updated as new runs arrive.