arXiv:2509.04830eess.AScs.SD2025-09被引 3

分析多语言语音合成质量在模型各层的编码方式,揭示早期自监督特征与音质相关。

Layer-wise Analysis for Quality of Multilingual Synthesized Speech

  • 通过参考模型逐层分析多语言语音模型中的质量信息编码
  • 早期自监督学习层特征与人工评分有显著相关性
  • 使用匹配参考数据对结果至关重要,适合语音质量评估研究者

尽管监督式语音质量预测器与人工评分有强相关性,但其依赖领域内标注数据,限制了跨领域的泛化能力。基于预训练自监督学习(SSL)和自动语音识别(ASR)模型的无监督方法是可行替代,但对其如何编码语音质量信息仍了解甚少。为深入理解多语言场景下语音质量各方面的编码机制,我们基于参考模型开展分层分析。结果表明,早期SSL层提取的特征与合成语音的人工评分存在相关性,而后期ASR层可预测非神经语音系统的质量及可懂度。同时,我们验证了使用匹配参考数据的重要性。

原文摘要 · Abstract (English)

While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised approaches based on pretrained self-supervised learning (SSL) based models and automatic speech recognition (ASR) models are a promising alternative; however, little is known about how these models encode information about speech quality. Towards the goal of better understanding how different aspects of speech quality are encoded in a multilingual setting, we present a layer-wise analysis of multilingual pretrained speech models based on reference modeling. We find that features extracted from early SSL layers show correlations with human ratings of synthesized speech, and later layers of ASR models can predict quality of non-neural systems as well as intelligibility. We also demonstrate the importance of using well-matched reference data.

语音合成质量评估多语言自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。