← Leaderboard

Nemotron 3 Nano 30B-A3B 4bit MLX

Tested Rank #13 of 33 · 5/8 GAUNTLET progress
Generalist — 79.5 / 100 G Agentic — 92 / 100 A Understanding — pending U Needle — 72.8 / 100 N Thinking — 96.7 / 100 T Live — pending L Engineering — pending E Throughput — 53.1 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
79.5
A
92
U
N
72.8
T
96.7
L
E
T
53.1

Specification

Parameters
30B total / ~3B active per token
Architecture
nemotron3_nano (hybrid Mamba2/attention MoE)
Size on disk
17.79 GB
Quantization
4bit
Format
MLX
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M14
Mean speed
86.7 tok/s across suites
Stall census
1 stall in 30 observed tests (3.3%)
Reasoning appetite
1,413 tokens mean · 16,383 max
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 206.7 / 260 15.9 13/13 132.4
Agentic tool-calling & protocol adherence 1b 147.2 / 160 18.4 8/8 107.7
Long-context retrieval & synthesis 1g 118.2 / 120 19.7 6/6 60.7
Long-context multi-needle (MRCR) 1g2 12.9 / 60 4.3 3/3 45.8

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.