arXiv:2606.31729eess.AScs.LG2026-06中稿 · Interspeech 26'

TTS评价不能只看自然度,还得看是否符合使用场景。

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation

论文配图:Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation
图 1 · 摘自论文原文
  • 在5种不同场景下测试SOTA TTS系统的表现差异
  • 自然度与适合度独立变化,优化一种可能损害另一种
  • 现有统一评价指标在表达性场景中存在盲区

文本转语音(TTS)评估仍是开放挑战。尽管早期目标是追求语音自然度,但近年来合成质量提升后,研究重点转向语音是否适合其使用场景——即“合适性”。本文考察了不同下游用途下人类感知的变化,测量了五种SOTA TTS系统在五个领域(AI助手、朗读、演员、动画角色、即兴说话者)中的合适性与类人程度。结果表明,合适性在不同领域间独立于自然度而变化;系统在朗读任务中表现优异,但在表达性任务中仍具挑战性,且针对某一领域的优化可能损害其他领域表现。此外,自然度评分往往惩罚风格化语音,却奖励即兴感。研究还揭示了通用评价指标在更富表现力场景中的盲点。结论表明,TTS性能并未“解决”,而是高度依赖目标领域,需采用情境感知的评估方式。

原文摘要 · Abstract (English)

Text-to-speech (TTS) evaluation is an open challenge. While the primary target was "naturalness," recent fidelity gains shifted focus toward "appropriateness" and whether speech is correct for its context. In this work, we examine how perception changes when the expected downstream use varies. We measure the appropriateness and human-likeness of five SOTA TTS systems across five domains: AI assistant, reader, actor, animated character, and spontaneous speaker. Results show appropriateness varies across domains independently of naturalness. While systems shine at reading, expressive domains remain challenging, and optimizing for one can degrade others. Furthermore, naturalness scores tend to penalize stylized speech while rewarding spontaneity. Finally, our study also highlights blind spots in one-size-fits-all evaluation metrics across more expressive domains. We demonstrate that TTS performance is not "solved" but depends on the target domain, requiring context-aware evaluation.

语音合成评估方法场景适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。