Model benchmarks

qwen3.6-14b-a3b-fablevibes local LLM performance

As of August 2026, qwen3.6-14b-a3b-fablevibes runs at up to 95.3 tok/s for local inference (best of 3 community benchmark runs across 2 GPUs).

LM Studio
ShareRedditX

Model size

14B

Peak speed

95.3 tok/s

Average speed

71.4 tok/s

Min memory

8.7 GB

Max context

5,483 tokens

Best quality

77.6

Benchmark runs

3

GPUs tested

2

Performance by hardware and tool

Every hardware/tool/quantization combination qwen3.6-14b-a3b-fablevibes has been benchmarked on, ranked by peak token generation speed. Last updated August 2026.

HardwareToolQuantBest tok/sAvg tok/sMemoryContextQualityRuns
AMD Radeon RX 6800 XTLM Studio95.3 tok/s84.4 tok/s8.7 GB5,483 tokens40.72
NVIDIA GeForce RTX 4070 Ti SUPERLM Studio45.2 tok/s45.2 tok/s14.5 GB5,127 tokens60.41

Frequently asked questions

How fast is qwen3.6-14b-a3b-fablevibes for local inference?
Across 3 community benchmark runs, qwen3.6-14b-a3b-fablevibes reaches up to 95.3 tok/s and averages 71.4 tok/s, with the fastest results on AMD Radeon RX 6800 XT.
How much memory does qwen3.6-14b-a3b-fablevibes need?
The leanest observed configuration used about 8.7 GB of memory.
Which tools have been used to run qwen3.6-14b-a3b-fablevibes?
Benchmarks were submitted using LM Studio. Results are community-contributed and updated as new runs arrive.