用听觉感知原理改进语音增强损失函数,让模型更关注人耳敏感的高频信息。
Loud-loss: A Perceptually Motivated Loss Function for Speech Enhancement Based on Equal-Loudness Contours
- 基于等响度曲线设计感知加权损失,按频率分配不同惩罚权重。
- 在VoiceBank+DEMAND数据集上,WB-PESQ评分从2.17提升至2.93。
- 方法通用性强,可适配各类语音增强模型,适合追求听感质量的研究者。
均方误差(MSE)是语音增强中广泛使用的损失函数,但其缺陷在于无法反映听觉感知质量。这是因为MSE会过度强调能量较高的低频成分,导致对人耳敏感的高频信息建模不足。为克服这一问题,本文提出一种基于心理声学原理的感知加权损失函数。该方法利用等响度曲线为重建误差分配频率相关的权重,使惩罚机制更符合人类听觉敏感度。所提损失函数具有模型无关性与灵活性,通用性强。在VoiceBank+DEMAND数据集上的实验表明,将GTCRN模型中的MSE替换为该损失后,WB-PESQ得分从2.17提升至2.93,显著改善了感知质量。
原文摘要 · Abstract (English)
The mean squared error (MSE) is a ubiquitous loss function for speech enhancement, but its problem is that the error cannot reflect the auditory perception quality. This is because MSE causes models to over-emphasize low-frequency components which has high energy, leading to the inadequate modeling of perceptually important high-frequency information. To overcome this limitation, we propose a perceptually-weighted loss function grounded in psychoacoustic principles. Specifically, it leverages equal-loudness contours to assign frequency-dependent weights to the reconstruction error, thereby penalizing deviations in a way aligning with human auditory sensitivity. The proposed loss is model-agnostic and flexible, demonstrating strong generality. Experiments on the VoiceBank+DEMAND dataset show that replacing MSE with our loss in a GTCRN model elevates the WB-PESQ score from 2.17 to 2.93-a significant improvement in perceptual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。