Benchmark result

Qwen3.8-27B-4bit on Apple M5 Max — 30.5 tok/s

Measured with oMLX 0.6.0.dev1 on August 14, 2026.

ShareRedditX

LLM / Model

Model Size
27B
Architecture
qwen
Memory Kind
model
Tool
oMLX v0.6.0.dev1
Scenarios

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
Overall Quality Score
86.85
coherence
91
decomposition
88
error_handling
80
tool_selection
84

The plan is logically structured and follows the required four-step flow, with clear dependencies from source gathering to reporting. Tool choices are mostly appropriate, though send_email is unnecessary and read_file is somewhat artificial. The task is well decomposed into sourcing, normalization, validation, and report generation. Error handling is solid, with a reasonable fallback strategy for failed sources and missing files, though some contingencies remain somewhat generic.

Code Generationcoding_agentCOMPLETED▶ Play Game
Overall Quality Score
81.30
correctness
78
performance
82
code_quality
84
completeness
85

Playable and fairly polished Breakout implementation with mobile support, scoring, lives, restart, and canvas rendering. It largely meets the requirements, but there are spec mismatches around HUD/brick overlap and some collision edge cases that reduce strict correctness.

Role Play & NarrativeroleplayCOMPLETED
Overall Quality Score
93.75
dialogue
90
immersion
96
consistency
94
narrative_arc
93

A highly effective in-character response with vivid tavern atmosphere, believable guarded warmth, and a strong emotional core. The main limitation is that it condenses the full five-turn arc into one sequence, but the character work is strong and consistent.

Research & AnalysisresearchCOMPLETED
Overall Quality Score
87.25
depth
92
clarity
90
insight
89
interpretation
78

Strong structure and generally solid comparative analysis, with good use of score gaps and task-by-task recommendations. However, the scaling interpretation is only moderately reliable because the response mixes log-linear fitting with a failed raw extrapolation, and it does not sufficiently separate parameter scaling from family/training differences.

GPU Acceleration

GPU Model
Apple M5 Max
VRAM
64 GB

Performance Metrics

Generation Speed
30.5
tokens/sec
TTFT
1273
milliseconds
Memory Usage
18.46
GB
Memory (Raw)
18904
MB

Context Information

Tokens Used
14112
Context Window
65,536

CPU & Memory

CPU Model
Apple M5 Max
CPU Threads
18
Physical Cores
18
CPU Frequency
N/A
Total RAM
64 GB

Operating System

OS Name
macos
OS Version
26.5.2
Kernel Version
25.5.0

Submission Details

Submitted:
8/14/2026, 8:54:49 PM
Client:
v0.4.36+97
Benchmark Result ID:
cmstfe5yk001s01p4kczrr9hz