arXiv:2409.08702eess.AScs.AI2024-09中稿 · publication in Com…被引 1

双分支并行网络统一处理语音增强与修复,效果更优且模型轻量。

A Dual-Branch Parallel Network for Speech Enhancement and Restoration

  • 采用双分支并行结构,分别负责噪声抑制与频谱重建。
  • 跨分支跳跃融合使抑制与生成能力互补,提升整体性能。
  • 在多种复杂噪声场景下表现优异,适合实际语音应用。

我们提出一种新型通用语音修复模型DBP-Net(双分支并行网络),有效应对真实世界中的复杂失真问题,包括噪声、混响和带宽压缩。不同于以往依赖单一处理路径或分离模型的方法,DBP-Net采用统一架构,包含两个并行分支:基于掩码的失真抑制分支和基于映射的频谱重建分支。其核心创新在于两分支间的参数共享与跨分支跳跃融合,即掩码分支的输出被显式融合至映射分支。该设计使DBP-Net在轻量框架内同时利用抑制与生成两种互补学习策略。实验表明,DBP-Net在综合性语音修复任务中显著优于现有基线方法,同时保持紧凑模型规模。结果表明,该模型为多样化失真场景下的统一语音增强与修复提供了高效且可扩展的解决方案。

原文摘要 · Abstract (English)

We present a novel general speech restoration model, DBP-Net (dual-branch parallel network), designed to effectively handle complex real-world distortions including noise, reverberation, and bandwidth degradation. Unlike prior approaches that rely on a single processing path or separate models for enhancement and restoration, DBP-Net introduces a unified architecture with dual parallel branches-a masking-based branch for distortion suppression and a mapping-based branch for spectrum reconstruction. A key innovation behind DBP-Net lies in the parameter sharing between the two branches and a cross-branch skip fusion, where the output of the masking branch is explicitly fused into the mapping branch. This design enables DBP-Net to simultaneously leverage complementary learning strategies-suppression and generation-within a lightweight framework. Experimental results show that DBP-Net significantly outperforms existing baselines in comprehensive speech restoration tasks while maintaining a compact model size. These findings suggest that DBP-Net offers an effective and scalable solution for unified speech enhancement and restoration in diverse distortion scenarios.

语音修复双分支网络轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。