arXiv:2510.26190cs.SDcs.CL2025-10被引 2

用问答评估语音合成的语义理解力,发现高准确率未必真听懂。

SP-MCQA: Evaluating Intelligibility of TTS Beyond the Word Level

  • 设计新闻片段问答任务,评估合成语音的关键信息保留度
  • 实验表明低错误率(WER)仍可能丢失关键信息,真实理解力不足
  • 适合关注语音合成实际可用性与评测标准升级的研究者

语音合成(TTS)的可理解性评估已陷入瓶颈,现有方法主要依赖词级准确率指标如字错误率(WER),无法反映真实语境下的语言理解需求。为此,本文提出一种新型主观评估方法——口语段落多选题问答(SP-MCQA),用于衡量合成语音中关键信息的准确性,并发布了一个时长8.76小时的新闻风格基准数据集SP-MCQA-Eval。实验结果表明,即使字错误率较低,关键信息准确率依然可能偏低,暴露出传统指标与实际理解能力之间的差距。SP-MCQA进一步揭示,当前最先进(SOTA)模型在文本归一化和语音准确性方面仍存在明显缺陷。该研究强调,在多数系统已达到高字错误率表现的当下,亟需建立更贴近人类认知、更高层次的评估标准。

原文摘要 · Abstract (English)

The evaluation of intelligibility for TTS has reached a bottleneck, as existing assessments heavily rely on word-by-word accuracy metrics such as WER, which fail to capture the complexity of real-world speech or reflect human comprehension needs. To address this, we propose Spoken-Passage Multiple-Choice Question Answering, a novel subjective approach evaluating the accuracy of key information in synthesized speech, and release SP-MCQA-Eval, an 8.76-hour news-style benchmark dataset for SP-MCQA evaluation. Our experiments reveal that low WER does not necessarily guarantee high key-information accuracy, exposing a gap between traditional metrics and practical intelligibility. SP-MCQA shows that even state-of-the-art (SOTA) models still lack robust text normalization and phonetic accuracy. This work underscores the urgent need for high-level, more life-like evaluation criteria now that many systems already excel at WER yet may fall short on real-world intelligibility.

语音合成可理解性评估多选题评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。