用GAN同时去混响并生成定向麦克风信号,提升复杂环境音质。
GAN-based Joint Dereverberation and Directional Filtering

- 基于GAN的联合去混响与定向滤波,一步完成信号重建。
- 相比串行处理方法,新模型在高阶定向信号上提升显著。
- 仅用输入输出信号即可估计方向图,适合实时语音处理。
近期神经定向滤波(NDF)可重建虚拟定向麦克风(VDM),精准还原多声源场景的空间线索。但在强混响环境下,空间线索难以分辨,限制了NDF的应用。本文提出神经去混响与定向滤波(NDDF)方法,实现去混响后的VDM信号重建。采用判别式训练与生成对抗网络(GAN)结合的模型,相较串行去混响与定向滤波基线表现更优。实验表明,当目标为高阶VDM时,GAN版本显著优于判别式变体。此外,本文提出一种仅依赖输入输出信号的方向图估计方法,适用于直接信号映射型空间滤波,无需显式滤波或掩码操作。
原文摘要 · Abstract (English)
Recently, neural directional filtering (NDF) enables reconstruction of a virtual directional microphone (VDM) with a desired directivity pattern, accurately rendering multi-source scenes by preserving spatial cues. In strongly reverberant environments, spatial cues become perceptually difficult to distinguish, limiting NDF-based spatial sound capture. This paper addresses this limitation with three contributions: First, we propose a neural dereverberation and directional filtering (NDDF) approach to reconstruct dereverberated VDM signals. Second, NDDF is implemented with discriminatively trained and generative adversarial network (GAN)-based models, compared with cascaded dereverberation and directional-filtering baselines. Experimental results indicate that the NDDF consistently surpasses the cascaded baselines. Additionally, the GAN-based NDDF outperforms the discriminative variant when addressing a high-order VDM target. Third, we introduce a method for directivity pattern estimation that relies solely on the input and output signals. This method is suitable for signal-mapping-based spatial filtering, which synthesizes the output signal directly without explicit filtering or masking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。