← Leaderboard

Qwen3-4B-2507 (non-thinking) 4bit MLX

Tested Rank #25 of 33 · 2/8 GAUNTLET progress
Generalist — 68.5 / 100 G Agentic — pending A Understanding — pending U Needle — pending N Thinking — pending T Live — pending L Engineering — pending E Throughput — 84.3 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
68.5
A
U
N
T
n/d
L
E
T
84.3

Specification

Parameters
4B
Architecture
qwen3
Size on disk
2.28 GB
Quantization
4bit
Format
MLX
Reasoning (CoT)
No
Internal ID
M18
Mean speed
137.6 tok/s across suites
Stall census
0 stalls in 14 observed tests (0.0%)
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 178.1 / 260 13.7 13/13 137.6

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.