← Leaderboard
Qwen3.8-27B Q6_K GGUF
Daily-driver trial RAN THE GAUNTLET Rank #1 of 33 · 8/8 GAUNTLET progress
G
99
A
100
U
90.1
N
89.3
T
96.5
L
75.7
E
98.5
T
8.1
Specification
- Parameters
- 27B dense
- Architecture
- qwen3_5
- Size on disk
- 23.36 GB
- Quantization
- Q6_K
- Format
- GGUF
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- M25b
- Mean speed
- 13.2 tok/s across suites
- Stall census
- 3 stalls in 86 observed tests (3.5%)
- Reasoning appetite
- 3,233 tokens mean · 31,137 max
- Model card
- huggingface.co
Suite results
| Suite | Score | Avg / 20 | Tests | tok/s |
|---|---|---|---|---|
| General capability (13-task real-workload suite) 1a | 257.4 / 260 | 19.8 | 13/13 | 18.2 |
| Agentic tool-calling & protocol adherence 1b | 160 / 160 | 20 | 8/8 | 18.3 |
| Coding depth 1c | 118.2 / 120 | 19.7 | 6/6 | 18.6 |
| Doc/OCR vision 1d | 155.2 / 160 | 19.4 | 8/8 | 11.7 |
| Doc/OCR — real-degraded tier 1d2 | 61 / 80 | 15.3 | 4/4 | 13.9 |
| Content-production depth 1e | 118.8 / 120 | 19.8 | 6/6 | 18 |
| Long-context retrieval & synthesis 1g | 118.8 / 120 | 19.8 | 6/6 | 4.4 |
| Long-context multi-needle (MRCR) 1g2 | 42 / 60 | 14 | 3/3 | 2.5 |
| Live one-shot builds (runtime-verified) 1h | 79 / 80 | 19.8 | 4/4 | — |
| Production replay (real agent workload) 1i | 163.2 / 240 | 13.6 | 12/12 | 13.4 |
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.