Benchmark result

gemma4:31b on Apple M5 Max — 7.5 tok/s

Measured with Ollama 0.23.2 on May 16, 2026.

ShareRedditX

LLM / Model

Model Size
31.3B
Quantization
Q4_K_M
Tool
Ollama v0.23.2

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
Overall Quality Score
89.25
coherence
91
decomposition
90
error_handling
86
tool_selection
88

The plan is well-structured and logically ordered, with clear dependencies and a sensible data-to-synthesis pipeline. Tool choices are mostly appropriate, though send_email is only useful if an external delivery is actually required. The task is broken into manageable steps and includes a reasonable backup strategy for weak quantitative sources, but error handling could be stronger by naming a specific fallback source and handling missing/heterogeneous benchmark formats more explicitly.

Code Generationcoding_agentCOMPLETED▶ Play Game
Overall Quality Score
69.40
correctness
62
performance
78
code_quality
74
completeness
68

A playable Breakout game is provided and many core features are present, including scoring, lives, restart, Canvas rendering, and basic mobile touch drag. However, several spec-critical requirements are only partially met or simplified, especially collision side detection, brick layout alignment, collision robustness, and full mobile event coverage.

Role Play & NarrativeroleplayCOMPLETED
Overall Quality Score
83.80
dialogue
84
immersion
92
consistency
88
narrative_arc
56

An evocative and believable opening that captures Aldwyn's guarded reaction and the tavern's mood very well, but it stops far short of the full narrative arc required by the prompt.

Research & AnalysisresearchCOMPLETED
Overall Quality Score
77.30
depth
82
clarity
84
insight
78
interpretation
68

Strong structure and generally clear recommendations, but the statistical interpretation is somewhat overconfident. The response captures broad diminishing-returns patterns and makes useful use-case guidance, yet it overstates the certainty of logarithmic scaling and does not fully justify extrapolation or cross-task comparisons.

GPU Acceleration

Processor Offload
100% GPU
GPU Model
Apple M5 Max
VRAM
64 GB
Driver
apple-unified:sysctl

Performance Metrics

Generation Speed
7.5
tokens/sec
TTFT
60036
milliseconds
Memory Usage
28.39
GB
Memory (Raw)
29068
MB

Context Information

Max context used
5.073
Context Window
65.536

CPU & Memory

CPU Model
Apple M5 Max
CPU Threads
18
Physical Cores
18
CPU Frequency
N/A
Total RAM
64 GB

Operating System

OS Name
macos
OS Version
26.5
Kernel Version
25.5.0

Submission Details

Submitted:
5/16/2026, 12:53:53 PM
Client:
v0.4.22+97
Benchmark Result ID:
cmp8cl0gr000101qtse9kgeur