Model benchmarks

zai-org/glm-4.6v-flash local LLM performance

As of August 2026, zai-org/glm-4.6v-flash runs at up to 54.5 tok/s for local inference (best of 2 community benchmark runs across 1 GPU).

LM Studio
ShareRedditX

Model size

Unknown

Peak speed

54.5 tok/s

Average speed

33.9 tok/s

Min memory

7.4 GB

Max context

7,853 tokens

Best quality

89.4

Benchmark runs

2

GPUs tested

1

Performance by hardware and tool

Every hardware/tool/quantization combination zai-org/glm-4.6v-flash has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
NVIDIA GeForce RTX 4070 Ti SUPERLM Studio54.5 tok/s54.5 tok/s11.0 GB7,853 tokens61.11
CPU onlyLM Studio13.2 tok/s13.2 tok/s7.4 GB1,282 tokens19.21

Frequently asked questions

How fast is zai-org/glm-4.6v-flash for local inference?
Across 2 community benchmark runs, zai-org/glm-4.6v-flash reaches up to 54.5 tok/s and averages 33.9 tok/s, with the fastest results on NVIDIA GeForce RTX 4070 Ti SUPER.
How much memory does zai-org/glm-4.6v-flash need?
The leanest observed configuration used about 7.4 GB of memory.
Which tools have been used to run zai-org/glm-4.6v-flash?
Benchmarks were submitted using LM Studio. Results are community-contributed and updated as new runs arrive.