在仅有100条标注数据下,用迁移学习预测语音质量,效果显著。
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
- 分两阶段迁移学习:先用自动标注噪声数据微调wav2vec 2.0,再适配挑战数据
- 在无参考条件下预测语音质量,BAK得分相关性达0.867,OVRL为0.711
- 通过人工降噪数据增强训练,显著提升SIG预测性能(从0.207升至0.516)
我们提出一种非侵入式语音质量预测系统,用于应对VoiceMOS 2024挑战赛第3赛道任务。该任务要求在无参考信号且仅提供100条主观标注语音样本的条件下,估计ITU-T P.835标准中的SIG、BAK和OVRL三项指标。系统采用wav2vec 2.0模型,结合两阶段迁移学习策略:首先在自动标注的噪声数据上进行初步微调,随后在挑战数据上进一步适应。官方评测中,系统在BAK预测上取得最佳表现(LCC=0.867),OVRL排名第二(LCC=0.711)。赛后实验表明,在第一阶段加入人工降噪数据后,SIG预测的相关性从0.207大幅提升至0.516。结果证明,在严重数据受限条件下,结合目标数据生成的迁移学习策略对预测P.835评分具有显著有效性。
原文摘要 · Abstract (English)
We present a system for non-intrusive prediction of speech quality in noisy and enhanced speech, developed for Track 3 of the VoiceMOS 2024 Challenge. The task required estimating the ITU-T P.835 metrics SIG, BAK, and OVRL without reference signals and with only 100 subjectively labeled utterances for training. Our approach uses wav2vec 2.0 with a two-stage transfer learning strategy: initial fine-tuning on automatically labeled noisy data, followed by adaptation to the challenge data. The system achieved the best performance on BAK prediction (LCC=0.867) and a very close second place in OVRL (LCC=0.711) in the official evaluation. Post-challenge experiments show that adding artificially degraded data to the first fine-tuning stage substantially improves SIG prediction, raising correlation with ground truth scores from 0.207 to 0.516. These results demonstrate that transfer learning with targeted data generation is effective for predicting P.835 scores under severe data constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。