语音质量预测挑战赛突破传统,多场景下自动评分效果显著提升。
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
- 采用检索与非自监督表征(如频谱图)提升评分精度
- 在少量标注数据下实现噪声、纯净及增强语音的高质量预测
- 涵盖语音合成、歌声合成等多场景,适合语音评估研究者
我们介绍了第三届 VoiceMOS 挑战赛,旨在推动人类语音主观评分的自动化预测研究。共设三个赛道:第一赛道针对语音合成系统生成的高保真语音样本质量预测;第二赛道覆盖多种系统、听众和语言的歌唱语音合成与语音转换样本评分;第三赛道为噪声、干净及增强语音的半监督质量预测,仅提供极少量标注训练数据。来自学术界与工业界的八支团队参与,多数表现优于基线系统。成功方法包括基于检索的方法以及使用频谱图、音高直方图等非自监督表示。结果表明该挑战显著推进了主观语音质量预测领域的发展。
原文摘要 · Abstract (English)
We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of ``zoomed-in'' high-quality samples from speech synthesis systems. The second track was to predict ratings of samples from singing voice synthesis and voice conversion with a large variety of systems, listeners, and languages. The third track was semi-supervised quality prediction for noisy, clean, and enhanced speech, where a very small amount of labeled training data was provided. Among the eight teams from both academia and industry, we found that many were able to outperform the baseline systems. Successful techniques included retrieval-based methods and the use of non-self-supervised representations like spectrograms and pitch histograms. These results showed that the challenge has advanced the field of subjective speech rating prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。