提出更可靠的语音评估指标TTSDS2,可精准衡量合成语音质量。
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
- 基于分布匹配设计新指标TTSDS2,提升评估鲁棒性。
- 在14种语言、多个领域中,相关性均超0.50,领先其他16项指标。
- 开源超1.1万条主观评分数据与持续更新的多语言评测基准。
文本转语音(TTS)系统的评估困难且成本高。主观指标如平均意见分(MOS)难以跨研究比较,客观指标虽常用却极少经过主观验证。近期的TTS系统已能生成与真实语音难以区分的合成语音,使现有评估方法面临挑战。本文提出文本转语音分布得分2(TTSDS2),是原版TTSDS的改进版本。在多种领域和语言中,它是16个对比指标中唯一一个在所有领域和主观评分上均达到斯皮尔曼相关系数高于0.50的指标。同时,我们发布了多项资源:包含超过11,000条主观评分的数据集;用于持续重建多语言测试集以避免数据泄露的流水线;以及涵盖14种语言的持续更新的TTS评估基准。
原文摘要 · Abstract (English)
Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but rarely validated against subjective ones. Both kinds of metrics are challenged by recent TTS systems capable of producing synthetic speech indistinguishable from real speech. In this work, we introduce Text to Speech Distribution Score 2 (TTSDS2), a more robust and improved version of TTSDS. Across a range of domains and languages, it is the only one out of 16 compared metrics to correlate with a Spearman correlation above 0.50 for every domain and subjective score evaluated. We also release a range of resources for evaluating synthetic speech close to real speech: A dataset with over 11,000 subjective opinion score ratings; a pipeline for continually recreating a multilingual test dataset to avoid data leakage; and a continually updated benchmark for TTS in 14 languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。