Model benchmarks

gemma4:26b-a4b-it-qat local LLM performance

As of July 2026, gemma4:26b-a4b-it-qat runs at up to 78.6 tok/s for local inference (best of 10 community benchmark runs across 2 GPUs).

OllamaQ4_0

Model size

25.2B

Peak speed

78.6 tok/s

Average speed

56.0 tok/s

Min memory

14.0 GB

Max context

6,441 tokens

Best quality

87.8

Benchmark runs

10

GPUs tested

2

Performance by hardware and tool

Every hardware/tool/quantization combination gemma4:26b-a4b-it-qat has been benchmarked on, ranked by peak token generation speed. Last updated July 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
CPU onlyOllamaQ4_078.6 tok/s67.9 tok/s14.0 GB6,203 tokens70.34
Apple M4 ProOllamaQ4_062.3 tok/s53.9 tok/s14.0 GB5,474 tokens70.25
NVIDIA GeForce RTX 4070 Laptop GPUOllamaQ4_019.3 tok/s19.3 tok/s14.1 GB6,441 tokens52.91

Frequently asked questions

How fast is gemma4:26b-a4b-it-qat for local inference?
Across 10 community benchmark runs, gemma4:26b-a4b-it-qat reaches up to 78.6 tok/s and averages 56.0 tok/s.
How much memory does gemma4:26b-a4b-it-qat need?
The leanest observed configuration used about 14.0 GB of memory (quantizations tested: Q4_0).
Which tools have been used to run gemma4:26b-a4b-it-qat?
Benchmarks were submitted using Ollama. Results are community-contributed and updated as new runs arrive.