arXiv:2603.29326cs.SDcs.AI2026-03

用分频段注意力模型实现低延迟语音降噪,兼顾音质与实时性。

Real-Time Band-Grouped Vocal Denoising Using Sigmoid-Driven Ideal Ratio Masking

  • 分频段编码解码结构结合频率注意力,降低处理延迟。
  • 总延迟低于10毫秒,宽带语音质量得分提升0.21(稳态噪声)。
  • 适合直播、语音通话等对实时性要求高的场景。

近年来,基于深度学习的实时语音降噪取得了显著进展,展现出人工智能在保持语音自然度的同时提升信噪比(SNR)的能力。然而,许多深度学习方法存在高延迟、需长时上下文的问题,难以适用于实时场景。为此,我们提出一种基于双曲正切驱动的理想比率掩码模型,采用谱损失训练以增强信噪比并最大化感知质量。该模型使用分频段编码解码架构与频率注意力机制,在总延迟低于10毫秒的前提下,于稳态噪声下实现了PESQ-WB提升0.21,非稳态噪声下提升0.12。

原文摘要 · Abstract (English)

Real-time, deep learning-based vocal denoising has seen significant progress over the past few years, demonstrating the capability of artificial intelligence in preserving the naturalness of the voice while increasing the signal-to-noise ratio (SNR). However, many deep learning approaches have high amounts of latency and require long frames of context, making them difficult to configure for live applications. To address these challenges, we propose a sigmoid-driven ideal ratio mask trained with a spectral loss to encourage an increased SNR and maximized perceptual quality of the voice. The proposed model uses a band-grouped encoder-decoder architecture with frequency attention and achieves a total latency of less than 10,ms, with PESQ-WB improvements of 0.21 on stationary noise and 0.12 on nonstationary noise.

语音降噪实时处理分频段注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。