← Leaderboard

Mistral Small 3.2 24B 8bit MLX

Keeper Rank #5 of 33 · 7/8 GAUNTLET progress
Generalist — 79 / 100 G Agentic — 92.5 / 100 A Understanding — 78.7 / 100 U Needle — 80 / 100 N Thinking — 100 / 100 T Live — pending L Engineering — 70 / 100 E Throughput — 7.9 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
79
A
92.5
U
78.7
N
80
T
100
L
E
70
T
7.9

Specification

Parameters
24B
Architecture
mistral3
Size on disk
25.93 GB
Quantization
8bit
Format
MLX
Reasoning (CoT)
No
Internal ID
M13
Mean speed
12.9 tok/s across suites
Stall census
0 stalls in 54 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 205.4 / 260 15.8 13/13 20.3
Agentic tool-calling & protocol adherence 1b 148 / 160 18.5 8/8 17.1
Coding depth 1c 84 / 120 14 6/6 14.8
Doc/OCR vision 1d 144 / 160 18 8/8 14
Doc/OCR — real-degraded tier 1d2 44.8 / 80 11.2 4/4 11.4
Content-production depth 1e 85.2 / 120 14.2 6/6 19.7
Long-context retrieval & synthesis 1g 117 / 120 19.5 6/6 2.5
Long-context multi-needle (MRCR) 1g2 27 / 60 9 3/3 3.6

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.