用可学习的语音质量评估模型提升语音增强效果,让听感更好。
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
- 用多指标语音质量评估模型做监督信号指导语音增强训练
- 在模拟和真实数据上均显著提升多个质量指标表现
- 适合缺乏干净参考语音的真实场景语音增强任务
语音质量评估(SQA)旨在预测语音信号在多种失真下的感知质量,与语音增强(SE)密切相关。尽管SQA模型广泛用于评估SE性能,但其在指导SE训练方面的潜力尚未充分挖掘。本文提出一种训练框架,利用在公开语音增强排行榜上训练的多指标SQA模型作为监督信号,指导语音增强模型训练。该方法克服了传统目标(如SI-SNR)与主观感知不一致、跨评估指标泛化能力差的问题,并支持在无干净参考信号的真实数据上训练。在模拟和真实测试集上的实验表明,基于SQA的训练能持续提升多种质量指标的表现。代码与模型检查点已开源。
原文摘要 · Abstract (English)
Speech quality assessment (SQA) aims to predict the perceived quality of speech signals under a wide range of distortions. It is inherently connected to speech enhancement (SE), which seeks to improve speech quality by removing unwanted signal components. While SQA models are widely used to evaluate SE performance, their potential to guide SE training remains underexplored. In this work, we investigate a training framework that leverages a SQA model, trained to predict multiple evaluation metrics from a public SE leaderboard, as a supervisory signal for SE. This approach addresses a key limitation of conventional SE objectives, such as SI-SNR, which often fail to align with perceptual quality and generalize poorly across evaluation metrics. Moreover, it enables training on real-world data where clean references are unavailable. Experiments on both simulated and real-world test sets show that SQA-guided training consistently improves performance across a range of quality metrics. Code and checkpoints are available at https://github.com/urgent-challenge/urgent2026_challenge_track2
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。