arXiv:2506.01023cs.SDcs.AI2025-06中稿 · Interspeech 2025被引 1

提出分阶段层级滤波框架,实时提升单通道语音增强效果

A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement

  • 分两阶段处理时频域信息,分离时序与频域滤波复杂度
  • 在相同资源下优于现有系统,显著提升语音清晰度
  • 适合实时语音增强场景,尤其对算力受限设备友好

本文提出一种融合子带处理与深度滤波的模型,充分利用目标时频(TF)单元及其邻近单元的信息,实现单通道语音增强。子带模块在输入端捕获邻近频域单元信息,深度滤波模块在输出端对目标及邻近TF单元进行联合滤波。为进一步提升性能,将深度滤波解耦为时序与频域分量,并引入两级框架,降低每阶段滤波系数预测复杂度。此外,提出TAConv模块以强化卷积特征提取能力。实验结果表明,所提层级深度滤波网络(HDF-Net)有效利用邻近时频信息,在资源消耗更少的情况下超越其他先进系统。

原文摘要 · Abstract (English)

This paper proposes a model that integrates sub-band processing and deep filtering to fully exploit information from the target time-frequency (TF) bin and its surrounding TF bins for single-channel speech enhancement. The sub-band module captures surrounding frequency bin information at the input, while the deep filtering module applies filtering at the output to both the target TF bin and its surrounding TF bins. To further improve the model performance, we decouple deep filtering into temporal and frequency components and introduce a two-stage framework, reducing the complexity of filter coefficient prediction at each stage. Additionally, we propose the TAConv module to strengthen convolutional feature extraction. Experimental results demonstrate that the proposed hierarchical deep filtering network (HDF-Net) effectively utilizes surrounding TF bin information and outperforms other advanced systems while using fewer resources.

语音增强深度学习实时处理滤波网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。