通过相对通道融合提升多通道语音增强效果
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
- 引入参考通道,让各通道与参考通道协同处理
- 在CHiME-3数据集上显著提升语音增强指标
- 适合需要空间信息的多麦克风语音系统
基于U-Net结构的神经多通道语音增强模型展现出良好性能和泛化能力。这类模型通常独立编码输入通道,并在网络后期才融合通道信息。本文提出一种新方法:从输入阶段即引入相对信息,将每个通道与参考通道并列堆叠处理。该策略利用通道间的相对差异,自适应地融合跨通道信息,从而捕获关键的空间特征,提升整体性能。在CHiME-3数据集上的实验表明,该方法在多种网络架构下均提升了语音增强指标。
原文摘要 · Abstract (English)
Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate the channels during later stages of the network. In this paper, we propose a novel modification of these models by incorporating relative information from the outset, where each channel is processed in conjunction with a reference channel through stacking. This input strategy exploits comparative differences to adaptively fuse information between channels, thereby capturing crucial spatial information and enhancing the overall performance. The experiments conducted on the CHiME-3 dataset demonstrate improvements in speech enhancement metrics across various architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。