Benchmark result

Ornith-1.5-35B-A3B-MLX-8bit on Apple M5 Max — 92.6 tok/s

Measured with oMLX 0.6.2 on August 19, 2026.

ShareRedditX

LLM / Model

Model Size
35B
Memory Kind
model
Tool
oMLX v0.6.2
Scenarios

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
Overall Quality Score
60.40
coherence
68
decomposition
79
error_handling
61
tool_selection
42

The response has a workable four-step structure and sensible dependency ordering, but it mixes planning with an ungrounded report and repeatedly acknowledges uncertainty instead of executing the stated workflow. Tool selection is weak because it relies on execute_code without clearly justifying actual data acquisition, and it incorrectly downplays the need for send_email only superficially. Error handling is present via backup ideas and cross-checking, but it remains generic and not fully operationalized.

Code Generationcoding_agentCOMPLETED▶ Play Game
Overall Quality Score
77.00
correctness
72
performance
82
code_quality
78
completeness
84

Overall, this is a solid single-file Breakout implementation with mobile support, scoring, lives, restart, and canvas rendering. It is mostly complete and playable, but it diverges from the exact startup behavior and includes non-solution text, which is a significant prompt-compliance issue.

Role Play & NarrativeroleplayCOMPLETED
Overall Quality Score
90.15
dialogue
89
immersion
91
consistency
93
narrative_arc
84

A strong, character-faithful opening that captures Aldwyn’s suspicion, grief, and hard-earned discipline with good sensory detail and natural dialogue. It is immersive and convincing, though it functions mainly as a first-turn response rather than a complete five-turn arc.

Research & AnalysisresearchCOMPLETED
Overall Quality Score
90.30
depth
92
clarity
94
insight
88
interpretation
90

Strong, well-structured analysis with correct arithmetic, a defensible logarithmic scaling conclusion, and clear trade-off-based recommendations. The main weakness is that it appropriately flags but does not fully resolve major confounds from mixing model families and unknown evaluation methodology.

GPU Acceleration

GPU Model
Apple M5 Max
VRAM
64 GB

Performance Metrics

Generation Speed
92.6
tokens/sec
TTFT
488
milliseconds
Memory Usage
36.89
GB
Memory (Raw)
37776
MB

Context Information

Tokens Used
26073
Context Window
65,536

CPU & Memory

CPU Model
Apple M5 Max
CPU Threads
18
Physical Cores
18
CPU Frequency
N/A
Total RAM
64 GB

Operating System

OS Name
macos
OS Version
26.5.2
Kernel Version
25.5.0

Submission Details

Submitted:
8/19/2026, 9:38:26 PM
Client:
v0.4.36+97
Benchmark Result ID:
cmt0m5j4o00fz01p4pne8a2jz