用自监督学习提升多采样率语音自然度评分预测性能
Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
- 在自监督模型中加入采样率无关卷积层,提取通用语音特征
- 在AMC 2025挑战赛中,一项指标排名第一,总排名第四
- 适合语音质量评估、跨采样率应用的研究者参考
我们提交至AudioMOS挑战赛(AMC)2025 Track 3的方案,针对多采样频率(SFs)语音的平均意见分(MOS)预测任务。所提模型将采样率无关(SFI)卷积层嵌入自监督学习(SSL)框架,实现对不同采样率语音的统一特征提取。为提升预测性能,采用从预训练非SFI-SSL模型中迁移知识,并利用大规模MOS数据集进行预训练。在挑战赛中,该方法在一项评价指标上排名第一,最终排名第四。我们还报告了消融实验结果,验证了模型关键设计的有效性。
原文摘要 · Abstract (English)
We introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer into a self-supervised learning (SSL) model to achieve SFI speech feature extraction for MOS prediction. We present some strategies to improve the MOS prediction performance of our model: distilling knowledge from a pretrained non-SFI-SSL model and pretraining with a large-scale MOS dataset. Our submission to the AMC 2025 Track 3 ranked the first in one evaluation metric and the fourth in the final ranking. We also report the results of our ablation study to investigate essential factors of our model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。