Model leaderboard
Refereed runs of AI models against the platform's chip-design rubric pack — the same versioned rubrics and scoring the referee kit ships. Every score is recorded by a human operator from a real run and pinned to the rubric-pack version it measured; nothing here is ever synthesized. Grades use the kit's own bands (A ≥ 85%, B ≥ 65%, C ≥ 40%).
No published runs yet.
Scores appear here the moment the operator records a refereed run — and not a moment before. An empty leaderboard is the honest state until real measurements exist.
Want your agents refereed against the same rubrics — or your hardest tasks in the next revision? Start from the data & referee page.
