Benchmark result

Qwen3.6-27B-MLX-4bit on Apple M5 Max — 32.1 tok/s

Measured with oMLX 0.31.3 on August 10, 2026.

ShareRedditX

LLM / Model

Model Size
27B
Architecture
qwen
Memory Kind
model
Tool
oMLX v0.31.3
Scenarios

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
Overall Quality Score
90.75
coherence
92
decomposition
94
error_handling
86
tool_selection
89

The plan is logically ordered and maps well to the report requirements. Tool choices are mostly appropriate, with search_web and execute_code used effectively; send_email is functionally unnecessary for the core task but acceptable if delivery is part of the workflow. The task is broken into clear, manageable steps that cover data collection, processing, report assembly, and delivery. Error handling is solid, though the fallback strategy could be more precise about validating alternative sources and preserving citation quality.

Code Generationcoding_agentCOMPLETED▶ Play Game
Overall Quality Score
74.00
correctness
72
performance
74
code_quality
78
completeness
70

A mostly playable Breakout implementation with good structure, styling, scoring, lives, restart flow, and mobile touch support. However, it diverges from several exact spec details, especially brick positioning, control semantics, and collision precision, so it is solid but not fully compliant.

Role Play & NarrativeroleplayCOMPLETED
Overall Quality Score
88.70
dialogue
90
immersion
94
consistency
92
narrative_arc
68

Highly convincing in voice, tension, and atmosphere, with a strong tell at the name 'Kestra.' However, it only establishes the opening probe and does not complete the required character progression across turns.

Research & AnalysisresearchCOMPLETED
Overall Quality Score
79.80
depth
84
clarity
88
insight
78
interpretation
74

Well-structured and clear, with good task-by-task comparisons and concrete recommendations. However, the scaling claims are somewhat overconfident relative to the limited data, and the extrapolation is not strongly justified statistically. Interpretation is mostly correct but methodologically under-supported.

GPU Acceleration

GPU Model
Apple M5 Max
VRAM
64 GB

Performance Metrics

Generation Speed
32.1
tokens/sec
TTFT
1072
milliseconds
Memory Usage
17.75
GB
Memory (Raw)
18180
MB

Context Information

Max context used
6395
Context Window
65,536

Throughput

Total Runtime
8m 58s

CPU & Memory

CPU Model
Apple M5 Max
CPU Threads
18
Physical Cores
18
CPU Frequency
N/A
Total RAM
64 GB

Operating System

OS Name
macos
OS Version
26.5.2
Kernel Version
25.5.0

Submission Details

Submitted:
8/10/2026, 6:22:31 PM
Client:
v0.4.35+97
Benchmark Result ID:
cmsnk6w98000001o2ncmse4l5