arXiv:2511.07185eess.AS2025-11被引 2

用深度学习让小麦克风阵列实现任意方向性音频捕捉

Neural Directional Filtering Using a Compact Microphone Array

  • 用神经网络从阵列信号生成复数掩码,模拟理想指向话筒
  • 在空间混叠频率以上仍保持方向性不变,支持复杂高阶模式
  • 可动态调整方向,且对未见场景有良好泛化能力

使用紧凑麦克风阵列实现期望的方向性模式在众多音频应用中至关重要。传统波束成形器的方向性受麦克风数量和阵列孔径限制,小型阵列性能通常下降。为此,我们提出神经方向滤波(NDF)方法,利用深度神经网络实现预设方向性模式的声场捕获。NDF从麦克风阵列信号计算单通道复数掩码,再应用于参考麦克风,生成逼近虚拟定向麦克风输出的结果。我们引入训练策略并设计数据依赖性度量以评估方向性模式与方向性因子。实验表明:(i) 即使在空间混叠频率以上仍保持频率无关的方向性;(ii) 可逼近多样且高阶方向性模式;(iii) 支持方向自由调节;(iv) 具备对未见条件的泛化能力。实验对比显示其优于传统波束成形与参数化方法。

原文摘要 · Abstract (English)

Beamforming with desired directivity patterns using compact microphone arrays is essential in many audio applications. Directivity patterns achievable using traditional beamformers depend on the number of microphones and the array aperture. Generally, their effectiveness degrades for compact arrays. To overcome these limitations, we propose a neural directional filtering (NDF) approach that leverages deep neural networks to enable sound capture with a predefined directivity pattern. The NDF computes a single-channel complex mask from the microphone array signals, which is then applied to a reference microphone to produce an output that approximates a virtual directional microphone with the desired directivity pattern. We introduce training strategies and propose data-dependent metrics to evaluate the directivity pattern and directivity factor. We show that the proposed method: i) achieves a frequency-invariant directivity pattern even above the spatial aliasing frequency, ii) can approximate diverse and higher-order patterns, iii) can steer the pattern in different directions, and iv) generalizes to unseen conditions. Lastly, experimental comparisons demonstrate superior performance over conventional beamforming and parametric approaches.

语音增强神经波束成形麦克风阵列方向性滤波

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。