Skip to content

Build a speaker-diverse real-speech evaluation set #6

Description

@Awshesh12

The entire real-speech claim rests on CS-FLEURS JA–EN read/test: 196 utterances, one speaker (SS), read speech, align-then-swap switch points. Three limitations compound:

  • One speaker — results are speaker-specific and we cannot separate "generalises" from "happens to suit this voice"
  • Read, not spontaneous — no disfluencies, repairs, or natural prosodic switch marking
  • Artificially dense switching — span-swapping produces more switches per utterance than natural speech, inflating absolute error rates

04-evaluation/eval-ft/docs/human-eval-protocol.md already specifies a human gold-set protocol that was never executed. This is the single highest-value thing we could add.

What's needed

  • Record 30–60 minutes of natural JA–EN bilingual speech from several speakers (meeting-like, spontaneous)
  • Transcribe under the script-policy conventions in the existing protocol
  • Add as a third frozen benchmark alongside CS-FLEURS and the FLEURS controls
  • Report all systems on it

Even 100 utterances from 5 speakers would materially strengthen every claim in the paper, and would let us test whether the CS-FLEURS ranking holds on spontaneous speech.

Related

Would also let us check whether the synthetic-to-real gap measured in §8 is larger on spontaneous speech than on read speech — a natural follow-up result.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

evaluationEvaluation harness, metrics, benchmarksresearchExploratory direction, new scope

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions