Benchmark result

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s

Measured with oMLX 0.31.3 on September 12, 2026.

ShareRedditX

LLM / Model

Model Size
4B
Memory Kind
model
Tool
oMLX v0.31.3

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
11.448 output tokens · 133.5 tok/s output · 1135.8 tok/s PP · 86.1s
Overall Quality Score
82.40
coherence
86
decomposition
92
error_handling
80
tool_selection
74

The plan is logically structured, uses exactly four steps, and cleanly maps to the report requirements. It decomposes the task well and includes a reasonable backup strategy. Tool choice is mostly appropriate, but read_file and execute_code are not strongly justified given the task can be completed largely via web sources and manual synthesis, and the response goes beyond planning by including detailed benchmark content.

Code Generationcoding_agentCOMPLETED▶ Play Game
33.151 output tokens · 121.3 tok/s output · 1986.2 tok/s PP · 273.6s
Overall Quality Score
84.60
correctness
84
performance
81
code_quality
86
completeness
90

A strong, mostly complete Breakout implementation with mobile controls, scoring, lives, restart flow, and canvas rendering. The main weaknesses are a few collision edge cases and a minor deliverable-format issue from the surrounding response text.

Role Play & NarrativeroleplayCOMPLETED
710 output tokens · 145.4 tok/s output · 1195.6 tok/s PP · 5.3s
Overall Quality Score
78.60
dialogue
79
immersion
86
consistency
82
narrative_arc
54

The response captures Aldwyn's haunted restraint, caution, and hidden grief very well, with vivid atmosphere and authentic voice. However, it only presents an opening beat rather than the full requested arc, so the narrative progression is incomplete.

Research & AnalysisresearchCOMPLETED
15.191 output tokens · 134.2 tok/s output · 1725.8 tok/s PP · 113.5s
Overall Quality Score
85.80
depth
86
clarity
88
insight
82
interpretation
90

Strong and mostly accurate analysis with clear task-by-task calculations, good comparative interpretation, and actionable recommendations. The main weaknesses are that the scaling claims rest on a very small, heterogeneous dataset and the extrapolation is necessarily uncertain.

GPU Acceleration

GPU Model
Apple M5 Max
VRAM
64 GB

Performance Metrics

Generation Speed
133.6
tokens/sec
PP
1510.8
tokens/sec
TTFT
381
milliseconds
Memory Usage
18.97
GB
Memory (Raw)
19426
MB

Context Information

Max context used
33.988
Context Window
262.144

Throughput

Total Output Tokens
60.500
Total Runtime
7m 59s

Inference Configuration

oMLX · per-model config · quant 4bit · engine 0.31.3
top_k
40
top_p
0.95
temperature
0.7
repetition_penalty
1
Engine features
reasoning effort
high
mtp
off
dflash
off
specprefill
off
turboquant_kv
off
thinking_budget
off

CPU & Memory

CPU Model
Apple M5 Max
CPU Threads
18
Physical Cores
18
CPU Frequency
N/A
Total RAM
64 GB

Operating System

OS Name
macos
OS Version
26.6.2
Kernel Version
25.6.0

Submission Details

Submitted:
9/12/2026, 10:27:24 AM
Client:
v0.4.59+97
Benchmark Result ID:
cmty8r0cz00hh01pda6qnswyv