arXiv:2604.10979eess.SPcs.SD2026-04

用深度学习实现降噪不损语音,尤其适合嘈杂环境下的语音通信。

Speech-preserving active noise control: a deep learning approach in reverberant environments

  • 基于卷积循环网络与谱结构映射,处理非线性混响环境中的噪声。
  • 在人群嘈杂等非平稳噪声下,降噪效果优于传统FxLMS算法。
  • 专设语音保留损失函数,保障语音自然度与可懂度,适合真实场景。

传统主动降噪(ANC)系统多基于FxLMS算法,依赖线性假设,在处理宽带非平稳噪声或非线性声学路径时表现受限,且常将语音与噪声一并消除,影响通信质量。本文提出一种语音保全型深度学习ANC系统,旨在复杂声学环境中实现稳定降噪同时有效保留语音。系统构建端到端控制架构,核心采用卷积循环网络(CRN),利用长短期记忆(LSTM)捕捉信号时序特征,并结合复谱映射(CSM)技术解决非线性失真问题。为在去噪同时保留语音,设计了专用语音保留损失函数,通过识别频谱结构特征,选择性保留目标语音。此外,为验证实际有效性,采用图像源法(ISM)构建高保真声学仿真环境,模拟真实混响效应。实验表明,所提深度ANC系统在非平稳噪声(如人群喧哗)下显著优于传统FxLMS算法;基于PESQ与STOI的评估结果证实,系统有效保持了语音的自然度与可懂度。

原文摘要 · Abstract (English)

Traditional Active Noise Control (ANC) systems are mostly based on FxLMS algorithms, but such algorithms rely on linear assumptions and are often limited in handling broadband non-stationary noise or nonlinear acoustic paths. Not only that, the traditional method is used to eliminating all signals together, and noise reduction often accidentally damages the voice signal and affects normal communication. To tackle these issues, this study proposes a speech preserving deep learning ANC system, which aims to achieve stable noise reduction while effectively retaining speech in a complex acoustic environment. This study builds an end-to-end control architecture, the core of which adopts a Convolutional Recurrent Network (CRN). The structure uses the long short-term memory (LSTM) network to capture the time-related characteristics of acoustic signals. Combined with complex spectrum mapping (CSM) technology, the nonlinear distortion problem is effectively solved. In order to retain useful voice while removing noise, this study also designs a special voice retention loss function. This design guidance model selectively retains the target voice while suppressing environmental noise by identifying the characteristics of the spectrum structure. In addition, in order to verify whether the system is effective in real scenes, we use the Image Source Method (ISM) to build a high-fidelity acoustic simulation environment, which also simulates the real reverberation effect. Experimental results demonstrate that the proposed Deep ANC system achieves significantly better noise reduction than the traditional FxLMS algorithm, especially for non-stationary noises like crowd babble. Meanwhile, PESQ and STOI based evaluations confirm that the system preserves both the naturalness and intelligibility of the target speech.

降噪语音保全深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。