Benchmark result

unsloth/Qwen3.8-Flash-Next-GGUF on NVIDIA GeForce RTX 5090 — 43.4 tok/s

Measured with vLLM on August 27, 2026.

ShareRedditX

LLM / Model

Model Size
Unknown
Architecture
qwen
Memory Kind
reserved
Tool
vLLM
Scenarios

LLM Quality Assessment

Agent WorkflowopenclawCOMPLETED
Overall Quality Score
93.20
coherence
94
decomposition
95
error_handling
91
tool_selection
92

Strong, well-structured 4-step plan with clear dependencies, appropriate use of search_web/read_file/execute_code, and a sensible fallback strategy. Minor concern: it relies on potentially hard-to-verify benchmark/speed sources and the claimed cache path may be speculative, but overall the planning is coherent and task-aligned.

Code Generationcoding_agentCOMPLETED▶ Play Game
Overall Quality Score
74.50
correctness
68
performance
74
code_quality
82
completeness
79

A polished and largely playable Breakout implementation with good visuals, fixed-timestep updates, scoring, lives, restart, and mobile input. However, a few spec mismatches and collision edge cases reduce strict correctness, especially around launch direction and robust AABB side resolution.

Role Play & NarrativeroleplayCOMPLETED
Overall Quality Score
93.95
dialogue
92
immersion
95
consistency
97
narrative_arc
88

A highly convincing Aldwyn voice: guarded, wary, and visibly shaken by Kestra's name while maintaining a mercenary edge. The sensory detail and internal tension are especially strong, though the excerpt stops before fully executing the later trust-and-choice beats required by the prompt.

Research & AnalysisresearchCOMPLETED
Overall Quality Score
87.30
depth
90
clarity
92
insight
86
interpretation
84

Strong, well-structured analysis with correct calculations, clear task-by-task comparisons, and practical recommendations. The main weakness is methodological: it infers scaling from heterogeneous model families without controlling for architecture, training, or benchmark conditions, so the scaling claims are directionally sound but not causal.

GPU Acceleration

GPU Model
NVIDIA GeForce RTX 5090
VRAM
32 GB
Driver
610.57.04

Performance Metrics

Generation Speed
43.4
tokens/sec
TTFT
2468
milliseconds
Reserved VRAM
30.89
GB
Memory (Raw)
31629
MB

Memory here is the VRAM pool vLLM pre-allocated (gpu_memory_utilization), not the size of the model weights. It is not comparable to the model memory Ollama, LM Studio, or oMLX report, so the efficiency figure reads lower for this run.

Context Information

Max context used
26149
Context Window
65,536

Throughput

Total Runtime
24m 33s

CPU & Memory

CPU Model
AMD Ryzen 9 7900X 12-Core Processor
CPU Threads
24
Physical Cores
12
CPU Frequency
N/A
Total RAM
62 GB

Operating System

OS Name
linux
OS Version
Fedora Linux 44 (KDE Plasma Desktop Edition)
Kernel Version
7.1.10-200.fc44.x86_64

Submission Details

Submitted:
8/27/2026, 2:42:31 PM
Client:
v0.4.43+97
Benchmark Result ID:
cmtbmth7r002001nojxqd82hz