arXiv:2502.01185cs.SDcs.AI2025-02被引 4

用新模型实现语音主动消除,比传统降噪效果更好。

Deep Active Speech Cancellation with Mamba-Masking Network

  • 提出Mamba-Masking结构,直接处理参考语音信号生成反向信号
  • 在语音场景下提升6.2dB,降噪场景下提升7.2dB,性能显著
  • 适合需要高精度语音抑制的实时通信与耳机场景

我们提出一种新型深度学习网络用于主动语音消除(ASC),超越传统主动降噪(ANC)方法,能有效消除噪声和语音信号。所提出的Mamba-Masking架构引入一种直接作用于编码参考信号的掩码机制,可在快速变化、高频的语音条件下实现自适应且精确对齐的反向信号生成。同时,采用多频段分割策略进一步提升各频带间的相位对齐效果。此外,设计了一种优化驱动的损失函数,为反向信号生成提供接近最优的监督信号。实验结果表明,该方法在ANC场景下性能提升达7.2dB,在ASC场景下提升达6.2dB,显著优于现有方法。

原文摘要 · Abstract (English)

We present a novel deep learning network for Active Speech Cancellation (ASC), advancing beyond Active Noise Cancellation (ANC) methods by effectively canceling both noise and speech signals. The proposed Mamba-Masking architecture introduces a masking mechanism that directly interacts with the encoded reference signal, enabling adaptive and precisely aligned anti-signal generation-even under rapidly changing, high-frequency conditions, as commonly found in speech. Complementing this, a multi-band segmentation strategy further improves phase alignment across frequency bands. Additionally, we introduce an optimization-driven loss function that provides near-optimal supervisory signals for anti-signal generation. Experimental results demonstrate substantial performance gains, achieving up to 7.2dB improvement in ANC scenarios and 6.2dB in ASC, significantly outperforming existing methods.

语音消除深度学习反向信号Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。