Skip to content

test(eval): publish a reproducible Terminal-Bench score #64

Description

@QiaolongLi1201

Goal

Replace readiness claims with a reproducible public Moss Terminal-Bench result when the official submission path is available.

Scope

  • Pin the benchmark revision, task manifest, Moss adapter revision, model, reasoning settings, and container image digest.
  • Run the existing 200-case harness preflight before any submission.
  • Record environment/tool versions, pass/fail/skip totals, cost and latency telemetry, and raw artifact checksums.
  • Keep credentials in environment/configured runners; never commit them or embed them in artifacts.
  • Publish the exact reproduction command and link the official leaderboard/submission result.

Acceptance criteria

  • A clean checkout can reproduce the submitted run from documented commands.
  • The official score, raw results, and Moss commit SHA agree.
  • Skips, retries, and infrastructure failures are reported separately from task failures.
  • The README/user guide makes no superiority claim beyond the measured result.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions