Goal
Replace readiness claims with a reproducible public Moss Terminal-Bench result when the official submission path is available.
Scope
- Pin the benchmark revision, task manifest, Moss adapter revision, model, reasoning settings, and container image digest.
- Run the existing 200-case harness preflight before any submission.
- Record environment/tool versions, pass/fail/skip totals, cost and latency telemetry, and raw artifact checksums.
- Keep credentials in environment/configured runners; never commit them or embed them in artifacts.
- Publish the exact reproduction command and link the official leaderboard/submission result.
Acceptance criteria
- A clean checkout can reproduce the submitted run from documented commands.
- The official score, raw results, and Moss commit SHA agree.
- Skips, retries, and infrastructure failures are reported separately from task failures.
- The README/user guide makes no superiority claim beyond the measured result.
Goal
Replace readiness claims with a reproducible public Moss Terminal-Bench result when the official submission path is available.
Scope
Acceptance criteria