arXiv:2411.12719cs.CLcs.LG2024-11中稿 · TMLR被引 7

改进MUSHRA评估法,让语音合成质量超越真人时也能公正打分。

Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation

  • 提出新版本MUSHRA,消除对真人参考语音的过度依赖。
  • 在印地语和泰米尔语上测试,492人参与,发现评分偏差和疲劳问题。
  • 发布24.6万条标注数据集MANGO,助力语音合成评估研究。

尽管文本转语音(TTS)模型快速进步,但一致且稳健的人类评估框架仍缺失。传统MOS测试难以区分相似模型,而CMOS的成对比较耗时。MUSHRA可同时评估多个TTS系统,但本文发现其依赖匹配真人参考语音,导致现代超真人质量的合成语音被不公平压低。我们通过492名听众在印地语和泰米尔语上的全面评估,识别出两大问题:(i) 参考匹配偏差,评分受真人参考影响过大;(ii) 判定模糊,因缺乏细粒度指导。为此提出两个改进变体:首个使高质量合成语音获得更公平评分;第二个降低评分波动,提升一致性。结合两者,实现更可靠、更精细的评估。同时发布MANGO数据集,含246,000条人类评分,为印度语言语音合成提供首个大规模偏好数据支持。

原文摘要 · Abstract (English)

Despite rapid advancements in TTS models, a consistent and robust human evaluation framework is still lacking. For example, MOS tests fail to differentiate between similar models, and CMOS's pairwise comparisons are time-intensive. The MUSHRA test is a promising alternative for evaluating multiple TTS systems simultaneously, but in this work we show that its reliance on matching human reference speech unduly penalises the scores of modern TTS systems that can exceed human speech quality. More specifically, we conduct a comprehensive assessment of the MUSHRA test, focusing on its sensitivity to factors such as rater variability, listener fatigue, and reference bias. Based on our extensive evaluation involving 492 human listeners across Hindi and Tamil we identify two primary shortcomings: (i) reference-matching bias, where raters are unduly influenced by the human reference, and (ii) judgement ambiguity, arising from a lack of clear fine-grained guidelines. To address these issues, we propose two refined variants of the MUSHRA test. The first variant enables fairer ratings for synthesized samples that surpass human reference quality. The second variant reduces ambiguity, as indicated by the relatively lower variance across raters. By combining these approaches, we achieve both more reliable and more fine-grained assessments. We also release MANGO, a massive dataset of 246,000 human ratings, the first-of-its-kind collection for Indian languages, aiding in analyzing human preferences and developing automatic metrics for evaluating TTS systems.

语音合成人类评估MUSHRA印度语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。