Benchmark result

Qwen3.8-Flash-Next-Dynamic-oQ5e-BF16-PLE-mtp on Apple M4 Max — 54.8 tok/s

Measured with oMLX 0.31.3 on September 14, 2026.

ShareRedditX

LLM / Model

Model Size
Unknown
Architecture
qwen
Memory Kind
model
Tool
oMLX v0.31.3

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
2.605 output tokens · 52.4 tok/s output · 184.9 tok/s PP · 51.9s
Overall Quality Score
76.40
coherence
78
decomposition
81
error_handling
70
tool_selection
74

The plan is logically structured and mostly aligned with the task, with clear dependencies and a sensible 2-parallel then 2-serial execution pattern. Tool choices are broadly appropriate, especially search_web for citations and execute_code for synthesis, but execute_code is somewhat overused for narrative generation and no read_file use is justified because no local data source is expected. Decomposition is strong and exactly meets the 4-step constraint, though Step 4 overlaps with Step 3 in final report assembly. Error handling is present and thoughtful, including backup sourcing and failure modes, but some mitigation details are speculative and cite potentially unreliable estimation rather than a robust fallback.

Code Generationcoding_agentCOMPLETED▶ Play Game
4.147 output tokens · 64.4 tok/s output · 220.3 tok/s PP · 68.2s
Overall Quality Score
75.90
correctness
72
performance
76
code_quality
78
completeness
85

A solid, playable single-file Breakout implementation with mobile input, scoring, lives, and win/lose states. However, it deviates from several strict spec details, especially brick placement and exact ball initialization, and its collision resolution is approximate rather than fully robust AABB side detection with last-collision tracking.

Role Play & NarrativeroleplayCOMPLETED
1.274 output tokens · 47.2 tok/s output · 200.8 tok/s PP · 29.2s
Overall Quality Score
82.15
dialogue
84
immersion
95
consistency
78
narrative_arc
58

An atmospheric, emotionally credible opening that captures Aldwyn's hidden wound and cautious temperament well, but it reads as a single scene rather than the required full arc of trust-building and commitment.

Research & AnalysisresearchCOMPLETED
5.247 output tokens · 55.0 tok/s output · 210.1 tok/s PP · 98.5s
Overall Quality Score
82.05
depth
88
clarity
90
insight
84
interpretation
69

Strong structure, clear task-by-task recommendations, and good use of the provided numbers. However, the core interpretation overstates the reliability of cross-family scaling conclusions and the extrapolation, since architectural and evaluation confounds are not controlled. The analysis is persuasive and actionable, but the statistical validity is only moderate.

GPU Acceleration

GPU Model
Apple M4 Max
VRAM
128 GB

Performance Metrics

Generation Speed
54.8
tokens/sec
PP
204.0
tokens/sec
TTFT
2825
milliseconds
Memory Usage
81.88
GB
Memory (Raw)
83848
MB

Context Information

Max context used
5.899
Context Window
262.144

Throughput

Total Output Tokens
13.273
Total Runtime
4m 48s

Inference Configuration

oMLX · per-model config · quant oQ5e · engine 0.31.3
min_p
0
top_k
20
top_p
0.95
temperature
1
repetition_penalty
1
Engine features
reasoning effort
medium
mtp
on
dflash
off
specprefill
off
turboquant_kv
off
thinking_budget
off

CPU & Memory

CPU Model
Apple M4 Max
CPU Threads
16
Physical Cores
16
CPU Frequency
N/A
Total RAM
128 GB

Operating System

OS Name
macos
OS Version
27.0
Kernel Version
27.0.0

Submission Details

Submitted:
9/14/2026, 5:26:05 AM
Client:
v0.4.59+97
Benchmark Result ID:
cmu0sv81k00ec01n41y9kwx4p