Benchmark result

Qwen3.8-Flash-Next-Dynamic-oQ5e-BF16-PLE-mtp on Apple M4 Max — 50.5 tok/s

Measured with oMLX 0.31.3 on September 14, 2026.

ShareRedditX

LLM / Model

Model Size
Unknown
Architecture
qwen
Memory Kind
model
Tool
oMLX v0.31.3

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
16.453 output tokens · 48.1 tok/s output · 186.9 tok/s PP · 344.5s
Overall Quality Score
89.80
coherence
92
decomposition
96
error_handling
90
tool_selection
84

Strong 4-step plan with clear dependencies, appropriate use of web search and code-based synthesis, and good backup handling. Minor issue: Step 4 is somewhat redundant with Step 3, and the plan risks stretching exact step boundaries by embedding report content inside the planning section, but overall it is well-structured and task-aligned.

Code Generationcoding_agentCOMPLETED▶ Play Game
19.762 output tokens · 57.9 tok/s output · 219.8 tok/s PP · 345.0s
Overall Quality Score
79.00
correctness
82
performance
74
code_quality
78
completeness
80

A mostly functional and polished Breakout implementation that meets most core requirements, including canvas rendering, mobile touch support, scoring, lives, restart, and win/lose states. However, the code is more elaborate than necessary, and a few collision/state-flow edge cases reduce confidence in strict spec adherence.

Role Play & NarrativeroleplayCOMPLETED
1.774 output tokens · 46.3 tok/s output · 188.6 tok/s PP · 40.9s
Overall Quality Score
90.55
dialogue
90
immersion
95
consistency
92
narrative_arc
78

Highly atmospheric and character-faithful, with excellent tension and a believable reaction to Kestra. The main weakness is structural: it establishes the emotional conflict well, but does not fully complete the requested multi-turn progression or final help/refuse decision.

Research & AnalysisresearchCOMPLETED
16.886 output tokens · 49.8 tok/s output · 200.6 tok/s PP · 342.8s
Overall Quality Score
88.00
depth
90
clarity
92
insight
88
interpretation
84

Strong, well-structured analysis with clear task-by-task comparisons, reasonable recommendations, and appropriate caution about confounding. The main weakness is methodological rigor: the scaling-law claims and 4T extrapolation are suggestive rather than statistically established because the data mix different model families and lack uncertainty estimates.

GPU Acceleration

GPU Model
Apple M4 Max
VRAM
128 GB

Performance Metrics

Generation Speed
50.5
tokens/sec
PP
198.9
tokens/sec
TTFT
3107
milliseconds
Memory Usage
82.28
GB
Memory (Raw)
84252
MB

Context Information

Max context used
20.641
Context Window
262.144

Throughput

Total Output Tokens
54.875
Total Runtime
18m 19s

Inference Configuration

oMLX · per-model config · quant oQ5e · engine 0.31.3
min_p
0
top_k
20
top_p
0.95
temperature
1
repetition_penalty
1
Engine features
reasoning effort
xhigh
mtp
on
dflash
off
specprefill
off
turboquant_kv
off
thinking_budget
off

CPU & Memory

CPU Model
Apple M4 Max
CPU Threads
16
Physical Cores
16
CPU Frequency
N/A
Total RAM
128 GB

Operating System

OS Name
macos
OS Version
27.0
Kernel Version
27.0.0

Submission Details

Submitted:
9/14/2026, 5:21:08 AM
Client:
v0.4.59+97
Benchmark Result ID:
cmu0soum900e501n40iolo1dh