arXiv:2411.11123cs.SDeess.AS2024-11被引 5

基于音高与频谱信息的声乐质量评估方法,显著提升预测准确性。

Pitch-and-Spectrum-Aware Singing Quality Assessment with Bias Correction and Model Fusion

  • 融合音高直方图与非量化神经编码器提取声乐特征
  • 在多个指标上超越所有竞品系统,表现最优
  • 适合关注声乐质量评估与模型优化的研究者

我们参与了 VoiceMOS Challenge 2024 的第 2 轨,目标是预测歌唱样本的平均意见得分(MOS)。我们的提交方案在所有参赛团队中(不含官方基线)排名第一。本文进一步改进该方案,提出一种新型的音高-频谱感知声乐质量评估方法(PS-SQA)。PS-SQA 基于自监督学习(SSL)的 MOS 预测器,分别利用音高直方图和非量化神经编码器提取歌唱音高与频谱信息。此外,该方法引入偏差校正策略,缓解低资源训练样本带来的预测偏差,并采用模型融合技术进一步提升精度。实验结果表明,所提 PS-SQA 在所有系统级指标上均显著优于其他竞争系统,验证了其强大的声乐质量评估能力。

原文摘要 · Abstract (English)

We participated in track 2 of the VoiceMOS Challenge 2024, which aimed to predict the mean opinion score (MOS) of singing samples. Our submission secured the first place among all participating teams, excluding the official baseline. In this paper, we further improve our submission and propose a novel Pitch-and-Spectrum-aware Singing Quality Assessment (PS-SQA) method. The PS-SQA is designed based on the self-supervised-learning (SSL) MOS predictor, incorporating singing pitch and spectral information, which are extracted using pitch histogram and non-quantized neural codec, respectively. Additionally, the PS-SQA introduces a bias correction strategy to address prediction biases caused by low-resource training samples, and employs model fusion technology to further enhance prediction accuracy. Experimental results confirm that our proposed PS-SQA significantly outperforms all competing systems across all system-level metrics, confirming its strong sing quality assessment capabilities.

声乐评估自监督学习模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。