提出新损失函数,让语音信号的幅度和相位自动配对,提升重建质量。
An Explicit Consistency-Preserving Loss Function for Phase Reconstruction and Speech Enhancement
- 直接约束幅度与相位构成真实信号的频谱,避免盲目估计相位。
- 在VB-DMD数据集上表现媲美传统方法,在WSJ0-CHiME3上显著优于相位显式估计法。
- 适合语音增强、相位恢复等需精确时频表示的任务,尤其在噪声环境有效。
本文提出一种新型一致性保持损失函数,用于在相位重建(PR)和语音增强(SE)任务中恢复相位信息。不同于传统方法直接用深度模型估计相位,本方法通过引入特定约束,直接生成一致的幅度-相位对。所提损失强制一组复数构成一个实信号的短时傅里叶变换(STFT)表示,即真实信号的频谱。该方法避免了原始相位估计的高不规则性和对时间偏移的敏感性。首先在相位重建任务上验证其可行性;随后在VB-DMD和WSJ0-CHiME3两个数据集上评估语音增强效果。在VB-DMD上性能接近传统方法;在更具挑战性的WSJ0-CHiME3数据集上,优于那些显式估计相位的方法。
原文摘要 · Abstract (English)
In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly estimate the phase using a deep model, our idea is to exploit ad-hoc constraints to directly generate a consistent pair of magnitude and phase. Specifically, the proposed loss forces a set of complex numbers to be a consistent short-time Fourier transform (STFT) representation, i.e., to be the spectrogram of a real signal. Our approach thus avoids the difficulty of estimating the original phase, which is highly unstructured and sensitive to time shift. The influence of our proposed loss is first assessed on a PR task, experimentally demonstrating that our approach is viable. Next, we show its effectiveness on an SE task, using both the VB-DMD and WSJ0-CHiME3 data sets. On VB-DMD, our approach is competitive with conventional solutions. On the challenging WSJ0-CHiME3 set, the proposed framework compares favourably over those techniques that explicitly estimate the phase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。