arXiv:2608.09425eess.AS2026-08中稿 · publication in IWA…

让声音定位模型通用化,适配各种麦克风阵列。

Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions

  • 用阵列传递函数+卷积网络处理多通道音频
  • 在混响和背景噪音下仍能准确定位方向
  • 可直接用于手机等未见过的设备

波达方向(DoA)估计是多通道音频处理的核心技术,但许多深度学习方法受限于训练时使用的麦克风阵列,难以泛化到新设备。本文提出一种基于实测或模拟复数方向阵列传递函数(ATF)的通用神经DoA估计框架,适用于真实世界的多麦克风设备。该方法对多通道频谱图与ATF元数据分别使用独立的卷积编码器,通过交叉注意力融合表征,并采用多源笛卡尔向量输出形式预测声源方向。在有混响和弥漫人声噪声的二维与三维定位任务中,实验表明所提方法能有效泛化至此前未见的阵列配置(包括类似手机的结构),性能下降不明显,且在性能上优于或媲美传统与学习型基线方法。

原文摘要 · Abstract (English)

Direction-of-arrival (DoA) estimation is a key component of multichannel audio processing, yet many deep learning approaches remain tied to the microphone arrays used during training and generalize poorly to unseen devices. This paper proposes an array-generic neural DoA estimation framework using measured or simulated complex directional array transfer functions (ATFs) matched to real-world multi-microphone devices. The method processes multichannel spectrograms and ATF metadata with separate convolutional encoders, fuses the resulting representations through cross-attention, and predicts source directions using a multi-source Cartesian vector output formulation. Experiments on simulated 2D and 3D localization tasks under reverberation and diffuse babble noise show that the proposed approach generalizes to previously unseen arrays, including mobile-phone-like configurations, without major performance degradation, while remaining competitive with conventional and learning-based baselines.

声源定位神经网络通用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。