arXiv:2506.02232eess.AScs.SD2025-06中稿 · INTERSPEECH 2025被引 1

用说话人模型提升歌声质量评分,效果显著优于传统方法。

Investigating the Reasonable Effectiveness of Speaker Pre-Trained Models and their Synergistic Power for SingMOS Prediction

  • 采用说话人识别预训练模型提取细微声学特征
  • 融合多模型后达到最新最优性能(无具体数值)
  • 适合语音质量评估与歌声合成研究者参考

本文聚焦于歌声主观评分(SingMOS)预测。先前研究已证明使用先进预训练模型(PTMs)可提升性能,但未深入探索说话人识别类预训练模型(如x-vector、ECAPA)。我们假设此类模型因具备说话人识别的预训练,能更有效捕捉合成歌声中的细微特征(如音高、音色、强度)。实验验证了该假设,对比了包括说话人模型与音乐类模型在内的多种SOTA PTMs。此外,提出一种新融合框架BATCH,利用巴氏距离实现多模型融合。通过BATCH融合说话人识别模型,本方法在所有单个模型及基线融合技术中表现最优,并刷新了当前最佳水平。

原文摘要 · Abstract (English)

In this study, we focus on Singing Voice Mean Opinion Score (SingMOS) prediction. Previous research have shown the performance benefit with the use of state-of-the-art (SOTA) pre-trained models (PTMs). However, they haven't explored speaker recognition speech PTMs (SPTMs) such as x-vector, ECAPA and we hypothesize that it will be the most effective for SingMOS prediction. We believe that due to their speaker recognition pre-training, it equips them to capture fine-grained vocal features (e.g., pitch, tone, intensity) from synthesized singing voices in a much more better way than other PTMs. Our experiments with SOTA PTMs including SPTMs and music PTMs validates the hypothesis. Additionally, we introduce a novel fusion framework, BATCH that uses Bhattacharya Distance for fusion of PTMs. Through BATCH with the fusion of speaker recognition SPTMs, we report the topmost performance comparison to all the individual PTMs and baseline fusion techniques as well as setting SOTA.

语音评估预训练模型歌声合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。