去掉混响语音的相位信息,反而能提升弱监督语音去混响效果。
Is Phase Really Needed for Weakly-Supervised Dereverberation ?
- 基于统计波场理论,分析混响相位在时频域的特性。
- 排除混响相位后,模型性能显著提升,证明相位信息无益。
- 适合研究弱监督语音处理或模型优化的研究者阅读。
在无监督或弱监督语音去混响方法中,训练时目标干净语音未知。此时,仅凭混响语音能否恢复有效信息成为关键问题。本文基于统计波场理论,揭示晚混响在时频域会引入白噪声式相位扰动,低频外相位基本无用。实验验证:在近期弱监督框架下,通过损失函数中剔除混响相位,模型性能明显提升,说明相位并非必要信息。
原文摘要 · Abstract (English)
In unsupervised or weakly-supervised approaches for speech dereverberation, the target clean (dry) signals are considered to be unknown during training. In that context, evaluating to what extent information can be retrieved from the sole knowledge of reverberant (wet) speech becomes critical. This work investigates the role of the reverberant (wet) phase in the time-frequency domain. Based on Statistical Wave Field Theory, we show that late reverberation perturbs phase components with white, uniformly distributed noise, except at low frequencies. Consequently, the wet phase carries limited useful information and is not essential for weakly supervised dereverberation. To validate this finding, we train dereverberation models under a recent weak supervision framework and demonstrate that performance can be significantly improved by excluding the reverberant phase from the loss function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。