arXiv:2606.21343eess.AScs.CL2026-06中稿 · Interspeech 2026

提出新评估框架,更好衡量语音重建中的可懂度与身份保留。

An Evaluation Framework for Text-to-Speech Voice Reconstruction

  • 用最优最差法评估可懂度和说话人身份感知
  • 发现传统指标对极难懂语音无效,引入双参考分布度量
  • 适用于语音障碍者语音重建系统评测,适合康复与语音合成研究者

使用文本转语音(TTS)进行语音重建为言语障碍者提供了一种沟通方式,旨在保留其说话人身份的同时提升可懂度。以往研究通常依赖平均意见得分(MOS)评估自然度和说话人相似性,但该方法敏感性与可靠性有限。本文提出一个包含主观与客观成分的评估框架:主观上采用带情境设定的最优最差法(BWS)评估可懂度与说话人身份感知;客观上发现标准度量无法预测高度不清晰说话人的重建效果,因此引入一种新的双参考分布度量,以评估可懂度与说话人身份之间的权衡。通过对17个零样本TTS系统在193名说话人上的输出进行评估,结果表明该框架提供了可靠且任务对齐的语音重建评估方法。

原文摘要 · Abstract (English)

Voice reconstruction using Text-to-Speech (TTS) offers a communication method for people with speech disorders, which aims to retain their speaker identity while improving intelligibility. Previous work generally relies on Mean Opinion Score (MOS) to evaluate naturalness and speaker similarity, but this has limited sensitivity and reliability. We propose an evaluation framework with subjective and objective components. Subjectively, we evaluate perceived intelligibility and speaker identity using Best Worst Scaling (BWS) with situational framing. Objectively, we demonstrate that standard measures fail to predict reconstruction success for highly unintelligible speakers, so we introduce a novel dual-reference distributional measure to assess the trade-off between intelligibility and speaker identity. By evaluating the output of 17 zero-shot TTS systems for 193 speakers, we show that our framework provides a reliable and task-aligned approach for assessing voice reconstruction.

语音重建TTS评估可懂度说话人身份

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。