Model benchmarks

google/gemma-4-12b local LLM performance

As of August 2026, google/gemma-4-12b runs at up to 32.4 tok/s for local inference (best of 2 community benchmark runs across 2 GPUs).

LM Studio
ShareRedditX

Model size

12B

Peak speed

32.4 tok/s

Average speed

28.7 tok/s

Min memory

9.3 GB

Max context

5,365 tokens

Best quality

85.6

Benchmark runs

2

GPUs tested

2

Performance by hardware and tool

Every hardware/tool/quantization combination google/gemma-4-12b has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
AMD Radeon RX 6800 XTLM Studio32.4 tok/s32.4 tok/s9.3 GB5,017 tokens73.81
NVIDIA GeForce RTX 4070 Ti SUPERLM Studio24.9 tok/s24.9 tok/s12.0 GB5,365 tokens75.41

Frequently asked questions

How fast is google/gemma-4-12b for local inference?
Across 2 community benchmark runs, google/gemma-4-12b reaches up to 32.4 tok/s and averages 28.7 tok/s, with the fastest results on AMD Radeon RX 6800 XT.
How much memory does google/gemma-4-12b need?
The leanest observed configuration used about 9.3 GB of memory.
Which tools have been used to run google/gemma-4-12b?
Benchmarks were submitted using LM Studio. Results are community-contributed and updated as new runs arrive.