用听感反馈优化语音增强,让生成更自然。
Aligning Generative Speech Enhancement with Perceptual Feedback
- 用人类听感代理评估模型输出,直接优化感知质量。
- 在深度噪声抑制挑战赛上提升56%的语音质量。
- 首次将偏好优化用于语音增强,适合听感研究者。
基于语言模型(LM)的语音增强近期成为有前景的方向,但现有方法主要依赖词元级似然目标,难以反映人类听觉感知。这种不匹配限制了进展,因为信号精度优化并不总能提升语音自然度或听感舒适度。本文提出一种感知对齐的LM语音增强方法,采用直接偏好优化(DPO)并以UTMOS(神经MOS预测器)作为人类评分的代理,直接引导模型生成更符合听感偏好的输出。该设计将训练目标与感知质量直接关联,可广泛应用于各类基于语言模型的语音增强框架。在2020年深度噪声抑制挑战赛测试集上,本方法持续提升语音质量指标,相对增益最高达56%。据我们所知,这是首个将感知反馈引入基于语言模型的语音增强工作,也是首个在语音增强领域应用DPO的研究,确立了一种新的感知对齐增强范式。
原文摘要 · Abstract (English)
Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives that weakly reflect human perception. This mismatch limits progress, as optimizing signal accuracy does not always improve naturalness or listening comfort. We address this gap by introducing a perceptually aligned LM-based SE approach. Our method applies Direct Preference Optimization (DPO) with UTMOS, a neural MOS predictor, as a proxy for human ratings, directly steering models toward perceptually preferred outputs. This design directly connects model training to perceptual quality and is broadly applicable within LM-based SE frameworks. On the Deep Noise Suppression Challenge 2020 test sets, our approach consistently improves speech quality metrics, achieving relative gains of up to 56%. To our knowledge, this is the first integration of perceptual feedback into LM-based SE and the first application of DPO in the SE domain, establishing a new paradigm for perceptually aligned enhancement with SE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。