arXiv:2502.21097cs.SDeess.AS2025-02

用生成对抗网络过滤麦克风阵列数据中的噪声和反射干扰

Deep learning-based filtering of cross-spectral matrices using generative adversarial networks

  • 设计GAN架构对固定尺寸的互谱矩阵进行转换
  • 在不同复杂度声压仿真数据上训练,实现五种变换任务
  • 适用于语音增强、声源定位等音频信号处理场景

本文提出一种基于深度学习的方法,用于过滤麦克风阵列数据中由环境噪声、反射或声源方向性带来的影响,这些数据以互谱矩阵形式表示。具体而言,我们设计了一种生成对抗网络(GAN)架构,用于处理固定尺寸的互谱矩阵。该模型使用为本研究定制的、具有不同复杂度的声压仿真数据进行训练。基于在自编码任务超参数优化中的结果,我们训练出一个优化模型,可执行五种由声压仿真中固有复杂度差异衍生出的不同转换任务。

原文摘要 · Abstract (English)

In this paper, we present a deep-learning method to filter out effects such as ambient noise, reflections, or source directivity from microphone array data represented as cross-spectral matrices. Specifically, we focus on a generative adversarial network (GAN) architecture designed to transform fixed-size cross-spectral matrices. Theses models were trained using sound pressure simulations of varying complexity developed for this purpose. Based on the results from applying these methods in a hyperparameter optimization of an auto-encoding task, we trained the optimized model to perform five distinct transformation tasks derived from different complexities inherent in our sound pressure simulations.

语音增强生成对抗网络音频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。