Contribute

The leaderboard is single-operator by design. The test surface doesn't have to be.

Phase 1 — now: ideas in, results ours

Every score on this site is run by one operator, on one controlled host, through one pipeline. That isn't a limitation we're working around — it's what makes the numbers comparable. Community-run results would mean uncontrolled hardware, uncontrolled configs, and a leaderboard measuring setups instead of models.

What we want from you right now:

Both go through GitHub issues. No forms, no queue theater — an issue tracker we actually read.

Phase 2 — planned: contributed test designs

Full test designs contributed by the community: you design it, we run it on the controlled host, you get credit on the resulting scores. Contributed tests go through the same contamination policy as everything else — prompts and rubrics stay private once they're live.

Community-submitted results

Maybe, eventually, as a clearly separated dataset. Never mixed into the verified leaderboard. A number we didn't measure is a number we won't publish as ours.