Model benchmarks
gemma4:e4b local LLM performance
As of August 2026, gemma4:e4b runs at up to 115.4 tok/s for local inference (best of 6 community benchmark runs across 4 GPUs).
OllamaQ4_K_M
Model size
8.0B
Peak speed
115.4 tok/s
Average speed
68.3 tok/s
Min memory
3.0 GB
Max context
131,072 tokens
Avg runtime / run
3m 3s
Avg quality
62.3
Benchmark runs
6
GPUs tested
4
Quality by task
Average LLM-judged quality (0–100) with the run-to-run spread shown as a P5–P95 band, overall and for each benchmark task, across all 6 runs. The low and high columns show how much the judge’s score varies between runs, and need at least two runs to display.
| Task | P5 (low) | Avg | P95 (high) |
|---|---|---|---|
| Overall | 58.3 | 62.3 | 67.3 |
| Agent Workflow | 61.5 | 71.9 | 82.0 |
| Code Generation | 38.9 | 44.6 | 50.8 |
| Role Play & Narrative | 58.8 | 74.1 | 88.5 |
| Research & Analysis | 51.9 | 58.8 | 68.1 |
Performance by hardware and tool
Every hardware/tool/quantization combination gemma4:e4b has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.
| Hardware | Tool | Quant | Best tok/s | Avg tok/s | Memory | Context | Quality | Runs |
|---|---|---|---|---|---|---|---|---|
| AMD Radeon RX 7900 XTX | Ollama | Q4_K_M | 115.4 tok/s | 115.4 tok/s | 10.9 GB | 65,536 tokens | 58.3 | 1 |
| CPU only | Ollama | Q4_K_M | 75.2 tok/s | 75.2 tok/s | 3.2 GB | 8,192 tokens | 61.8 | 1 |
| NVIDIA GeForce RTX 4070 Laptop GPU | Ollama | Q4_K_M | 57.6 tok/s | 57.6 tok/s | 3.1 GB | 8,192 tokens | 64.5 | 1 |
| Apple M4 Pro | Ollama | Q4_K_M | 55.4 tok/s | 54.8 tok/s | 3.1 GB | 131,072 tokens | 60.6 | 2 |
| NVIDIA GeForce RTX 4060 Laptop GPU | Ollama | Q4_K_M | 51.8 tok/s | 51.8 tok/s | 3.0 GB | 8,192 tokens | 68.3 | 1 |
Frequently asked questions
- Is gemma4:e4b good for coding?
- In our benchmarks, gemma4:e4b scores 44.6/100 for coding. It runs at about 68.3 tok/s, so if you want more speed, unsloth/Qwen3.8-27B-GGUF:IQ3_S is faster (~104.3 tok/s) and still scores well for coding (83.0/100).
- Is gemma4:e4b good for agentic (tool-using) tasks?
- In our benchmarks, gemma4:e4b scores 71.9/100 for agentic workflows. It runs at about 68.3 tok/s, so if you want more speed, Tiel-Coder-35B-A3B-Q4_K_S is faster (~165.0 tok/s) and still scores well for agentic workflows (89.3/100).
- How fast is gemma4:e4b for local inference?
- Across 6 community benchmark runs, gemma4:e4b reaches up to 115.4 tok/s and averages 68.3 tok/s, with the fastest results on AMD Radeon RX 7900 XTX.
- How much memory does gemma4:e4b need?
- The leanest observed configuration used about 3.0 GB of memory (quantizations tested: Q4_K_M).
- Which tools have been used to run gemma4:e4b?
- Benchmarks were submitted using Ollama. Results are community-contributed and updated as new runs arrive.