Five trials. One exact system.
Every leaderboard entry represents an exact model-and-agent combination. Each system runs five valid trials per task. With 20 tasks, an official V1 entry requires 100 valid trials. Strict pass remains the authoritative binary outcome beside the multidimensional score.
Trust gates come before points.
Qualification checks decide whether a task or result is valid. They do not award extra model points.
Recovery Quality
Each dimension is normalized from 0 to 1. The weighted base quality is calculated before safety docking.
Unsafe recovery cannot score well.
ColdStart Recovery Score
Recovery Quality averages post-docking scores across valid trials. Reliability reflects strict-pass consistency. Efficiency uses cost and runtime only for successful trials, so cheap failures never gain an advantage.
Measured after calibration.
Difficulty is based on the highest strict-pass rate achieved by selected calibration systems, then frozen at V1 release: Hard 0–20%; Medium above 20–60%; Easy above 60–80%. Anything above 80% is excluded as insufficiently difficult. The final corpus target is 14 Hard, 4 Medium and 2 Easy tasks.
Failures need the right label.
Incorrect repair, timeout and inability to solve are valid scored model outcomes. Infrastructure failures are excluded and retried. A defective verifier, leaked task or unfair condition invalidates the result without penalizing the model. Every attempt remains preserved for audit.