基于可解释深度网络的多通道波束成形器,无需说话人信息即可提升语音增强效果。
Explainable DNN-based Beamformer with Postfilter
- 两阶段设计:先用固定权重构建波束成形器,再用时变后滤波器优化输出。
- 在真实场景下显著提升语音清晰度,且无需提前知道说话人位置或活动状态。
- 引入注意力机制并分析空间信息利用方式,提升模型可解释性,适合语音处理研究者。
本文提出一种可解释的基于深度神经网络的波束成形器与后滤波器结合方法(ExNet-BF+PF),用于多通道信号处理。该方法采用两阶段处理流程:第一阶段使用时不变权重构建多通道空间滤波器(即波束成形器);第二阶段在波束成形输出端应用时变单通道后滤波器。此外,借鉴噪声和混响环境中表现优异的注意力机制,进一步提升语音增强性能。本研究还填补了现有文献空白,通过深入的空间分析揭示网络在处理过程中如何利用空间信息,从而增进对模型功能的理解。实验结果表明,该方法训练简单、性能优越,且无需依赖说话人活动先验知识。
原文摘要 · Abstract (English)
This paper introduces an explainable DNN-based beamformer with a postfilter (ExNet-BF+PF) for multichannel signal processing. Our approach combines the U-Net network with a beamformer structure to address this problem. The method involves a two-stage processing pipeline. In the first stage, time-invariant weights are applied to construct a multichannel spatial filter, namely a beamformer. In the second stage, a time-varying single-channel post-filter is applied at the beamformer output. Additionally, we incorporate an attention mechanism inspired by its successful application in noisy and reverberant environments to improve speech enhancement further. Furthermore, our study fills a gap in the existing literature by conducting a thorough spatial analysis of the network's performance. Specifically, we examine how the network utilizes spatial information during processing. This analysis yields valuable insights into the network's functionality, thereby enhancing our understanding of its overall performance. Experimental results demonstrate that our approach is not only straightforward to train but also yields superior results, obviating the necessity for prior knowledge of the speaker's activity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。