提出自适应频域损失,改善语音增强中高频段失真问题。
Towards Balanced Spectral Reconstruction: Spectrally Adaptive Loss for Streaming Speech Enhancement

- 设计平滑的频域加权损失,缓解相位补偿导致的中高频衰减。
- 信号依赖型自适应损失在中频区提升重建效果,全频段更均衡。
- 适用于低延迟流式语音增强场景,适合实时应用开发人员。
本文针对轻量级流式语音增强中因幅度-相位补偿效应导致的中高频段幅度过度衰减问题,提出了两种谱加权STFT损失函数。其中,Sigmoid加权损失对相位感知贡献施加平滑的频率依赖调制;信号依赖型谱自适应损失进一步将调制条件基于真实对数幅度谱。为评估所提目标,额外设计了HyST-Net,一种融合混合MHA-GRU的轻量级骨干网络,适用于低延迟流式场景。实验结果表明,两种损失均在高频谱重建上取得一致提升,而谱自适应损失还进一步改善了中频区域,实现全频段更均衡的谱重构。
原文摘要 · Abstract (English)
This paper proposes two spectrally weighted STFT loss functions for lightweight streaming speech enhancement, addressing the magnitude over-attenuation in mid-to-high frequency regions caused by the magnitude-phase compensation effect. The proposed sigmoid-weighted loss applies a smooth frequency-dependent modulation to the phase-aware contribution, while the signal-dependent spectrally adaptive loss further conditions the modulation on the ground-truth log-magnitude spectrogram. To evaluate the proposed objectives, we additionally design HyST-Net, a lightweight and competitive backbone with hybrid MHA-GRU spectral-temporal modelling for low-latency streaming scenarios. Experimental results exhibit consistent improvements in high-frequency spectral reconstruction for both losses. The spectrally adaptive loss further enhances the mid-frequency region, resulting in a more balanced spectral reconstruction across the full frequency range.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。