Skip to content

Evaluation roadmap v2 #96

Description

@dangng2004

This supersedes the v1 roadmap #51 as it has led to this paper: acl_latex.pdf.

That benchmark measures systems' error detection ability, which is good for a start, but we also need a metric for the quality of each comment and the overall feedback.

Some desiderata of the overall feedback:

  • high-level feedback that is less standard and includes more substance.
  • Prioritization of the issues: do more serious issues get foregrounded?
  • Clarity: is the communication clear?
  • Balance: does it balance the strengths and weaknesses of a paper?
  • Conciseness: does it communicate in a concise manner?

Some desiderata for the comments:

  • Relevance
  • Accuracy

We will likely use LLM as a judge for this phase of the benchmark.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    evaluationEvaluation related issueshelp wantedExtra attention is needed

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions