arXiv:2602.08556cs.SD2026-02中稿 · IEEE Transactions …

让语音增强模型更好理解相位的周期性,提升降噪效果

Global Rotation Equivariant Phase Modeling for Speech Enhancement with Deep Magnitude-Phase Interaction

  • 设计相位流保持圆形几何特性的双流网络
  • 相位重建任务中相位距离降低20%以上
  • 适合需要精确相位建模的语音增强场景

深度学习虽推动了语音增强(SE)发展,但有效建模相位仍具挑战,因传统网络在平坦欧氏空间运行,难以捕捉相位内在的环形拓扑结构。为此,本文提出一种幅度-相位双流框架,通过施加全局旋转等变性(GRE)特性,使相位流与内在圆形几何对齐。具体包括:基于幅度的信息交互卷积模块(MPICM),以及统一特征融合的混合注意力双前馈网络(HADF)瓶颈,二者均设计为保持相位流的GRE性质。在相位重建、降噪、去混响和带宽扩展任务上进行全面评估,结果表明该方法优于多个先进基线。尤其在零样本跨语料库降噪中,PESQ得分提升超过0.1;整体在包含多种失真的通用语音增强任务中也表现更优。定性分析显示,学习到的相位特征呈现明显的周期性模式,符合相位固有的环形本质。代码已开源。

原文摘要 · Abstract (English)

While deep learning has advanced speech enhancement (SE), effective phase modeling remains challenging, as conventional networks typically operate within a flat Euclidean feature space, which is not easy to model the underlying circular topology of the phase. To address this, we propose a magnitude-phase dual-stream framework that aligns the phase stream with its intrinsic circular geometry by enforcing Global Rotation Equivariance (GRE) characteristic. Specifically, we introduce a Magnitude-Phase Interactive Convolutional Module (MPICM) for modulus-based information exchange and a Hybrid-Attention Dual Feed-Forward Network (HADF) bottleneck for unified feature fusion, both of which are designed to preserve GRE in the phase stream. Comprehensive evaluations are conducted across phase retrieval, denoising, dereverberation, and bandwidth extension tasks to validate the superiority of the proposed method over multiple advanced baselines. Notably, the proposed architecture reduces Phase Distance by over 20\% in the phase retrieval task and improves PESQ by more than 0.1 in zero-shot cross-corpus denoising evaluations. The overall superiority is also established in universal SE tasks involving mixed distortions. Qualitative analysis further reveals that the learned phase features exhibit distinct periodic patterns, which are consistent with the intrinsic circular nature of the phase. The source code is available at https://github.com/wangchengzhong/GRE-Net.

语音增强相位建模等变网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。