Benchmark result

Qwen3.8-Flash-Next-Dynamic-oQ5e-BF16-PLE-mtp on Apple M4 Max — 49.2 tok/s

Measured with oMLX 0.31.3 on September 14, 2026.

ShareRedditX

LLM / Model

Model Size
Unknown
Architecture
qwen
Memory Kind
model
Tool
oMLX v0.31.3

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
8.260 output tokens · 47.3 tok/s output · 184.5 tok/s PP · 177.0s
Overall Quality Score
90.95
coherence
92
decomposition
94
error_handling
88
tool_selection
89

The plan is well-structured, follows the required 4-step format, and maintains clear dependencies from data collection to report generation. Tool choices are appropriate overall, though send_email is correctly omitted and the use of read_file as a backup is sensible but somewhat speculative without a guaranteed file path. The decomposition cleanly separates sourcing, backup validation, normalization, and synthesis. Error handling is strong, with a backup strategy and comparability concerns acknowledged, but the plan could be more explicit about how it will handle mismatched benchmark hardware and missing metrics.

Code Generationcoding_agentCOMPLETED▶ Play Game
26.266 output tokens · 54.6 tok/s output · 205.2 tok/s PP · 485.0s
Overall Quality Score
22.20
correctness
12
performance
34
code_quality
28
completeness
22

Visually elaborate but not a valid solution to the requested Breakout game. The script contains a fatal syntax error and deviates from several core gameplay requirements, so it is unlikely to run correctly or meet the spec.

Role Play & NarrativeroleplayCOMPLETED
4.598 output tokens · 44.0 tok/s output · 203.5 tok/s PP · 107.0s
Overall Quality Score
96.30
dialogue
95
immersion
97
consistency
98
narrative_arc
93

A highly convincing portrayal of Aldwyn: guarded, haunted, and cautious without being melodramatic. The sensory detail and subtext are excellent, and the dialogue feels lived-in and fitting. It succeeds especially well as an opening response, though it only partially fulfills the full multi-turn progression and end-state commitment specified in the prompt.

Research & AnalysisresearchCOMPLETED
21.801 output tokens · 51.1 tok/s output · 202.4 tok/s PP · 430.2s
Overall Quality Score
86.00
depth
92
clarity
91
insight
88
interpretation
74

Strong structure and clear task-by-task recommendations, with solid qualitative scaling judgments. However, the quantitative interpretation is weakened by confounded cross-model comparisons and an extrapolation that is plausible but not robust.

GPU Acceleration

GPU Model
Apple M4 Max
VRAM
128 GB

Performance Metrics

Generation Speed
49.2
tokens/sec
PP
198.9
tokens/sec
TTFT
3130
milliseconds
Memory Usage
81.77
GB
Memory (Raw)
83735
MB

Context Information

Max context used
27.145
Context Window
262.144

Throughput

Total Output Tokens
60.925
Total Runtime
20m 23s

Inference Configuration

oMLX · per-model config · quant oQ5e · engine 0.31.3
min_p
0
top_k
20
top_p
0.95
temperature
1
repetition_penalty
1
Engine features
mtp
on
dflash
off
specprefill
off
turboquant_kv
off
thinking_budget
off

CPU & Memory

CPU Model
Apple M4 Max
CPU Threads
16
Physical Cores
16
CPU Frequency
N/A
Total RAM
128 GB

Operating System

OS Name
macos
OS Version
27.0
Kernel Version
27.0.0

Submission Details

Submitted:
9/14/2026, 4:50:09 AM
Client:
v0.4.59+97
Benchmark Result ID:
cmu0rl0y200dy01n44tmtcp2n