Benchmark result

Qwen3.8-Flash-Next-Dynamic-oQ5e-BF16-PLE-mtp on Apple M4 Max — 55.0 tok/s

Measured with oMLX 0.31.3 on September 14, 2026.

ShareRedditX

LLM / Model

Model Size
Unknown
Architecture
qwen
Memory Kind
model
Tool
oMLX v0.31.3

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
3.213 output tokens · 54.0 tok/s output · 191.5 tok/s PP · 61.6s
Overall Quality Score
76.50
coherence
78
decomposition
82
error_handling
69
tool_selection
74

The plan is logically structured and stays within the 4-step limit, with clear dependencies and a sensible separation of data gathering, computation, and synthesis. Tool choice is mostly appropriate, especially using search_web for sources and execute_code for table assembly, but the response overcommits to specific benchmarks and sources without showing how exact numbers will be validated. Error handling is present and includes a backup strategy, though it is somewhat generic and relies on web results being available rather than specifying stronger contingency handling for missing or inconsistent metrics.

Code Generationcoding_agentCOMPLETED▶ Play Game
3.802 output tokens · 64.0 tok/s output · 220.0 tok/s PP · 63.3s
Overall Quality Score
73.80
correctness
72
performance
74
code_quality
78
completeness
68

A mostly functional single-file Breakout game with solid structure and canvas rendering, but it misses several exact spec details and has a few interaction/collision bugs that affect correctness and completeness.

Role Play & NarrativeroleplayCOMPLETED
1.311 output tokens · 48.3 tok/s output · 194.9 tok/s PP · 29.4s
Overall Quality Score
85.00
dialogue
84
immersion
92
consistency
86
narrative_arc
68

An evocative, character-faithful opening that captures Aldwyn's haunted restraint and tavern atmosphere very well, but it functions more like a strong first turn than a complete five-turn arc.

Research & AnalysisresearchCOMPLETED
5.042 output tokens · 54.0 tok/s output · 209.1 tok/s PP · 96.6s
Overall Quality Score
80.05
depth
88
clarity
90
insight
84
interpretation
61

The response is strong on structure, task-by-task comparison, and actionable recommendations, but its quantitative interpretation is only moderately reliable because it infers a scaling law from sparse, confounded data. It provides useful directional insight, yet several numerical and causal claims overreach the evidence.

GPU Acceleration

GPU Model
Apple M4 Max
VRAM
128 GB

Performance Metrics

Generation Speed
55.0
tokens/sec
PP
203.9
tokens/sec
TTFT
2828
milliseconds
Memory Usage
82.19
GB
Memory (Raw)
84165
MB

Context Information

Max context used
5.694
Context Window
262.144

Throughput

Total Output Tokens
13.368
Total Runtime
4m 45s

Inference Configuration

oMLX · per-model config · quant oQ5e · engine 0.31.3
min_p
0
top_k
20
top_p
0.95
temperature
0.6
repetition_penalty
1
Engine features
reasoning effort
medium
mtp
on
dflash
off
specprefill
off
turboquant_kv
off
thinking_budget
off

CPU & Memory

CPU Model
Apple M4 Max
CPU Threads
16
Physical Cores
16
CPU Frequency
N/A
Total RAM
128 GB

Operating System

OS Name
macos
OS Version
27.0
Kernel Version
27.0.0

Submission Details

Submitted:
9/14/2026, 5:31:01 AM
Client:
v0.4.59+97
Benchmark Result ID:
cmu0t1k3t00ej01n4l23caobn