arXiv:2507.07066cs.SDcs.AI2025-07被引 2

自监督模型LAM实现高效精准声源定位,兼顾可解释性与适应性。

Latent Acoustic Mapping for Direction of Arrival Estimation: A Self-Supervised Approach

  • 通过自监督学习生成高分辨率声场映射图,无需大量标注数据
  • 在LOCATA和STARSS上性能媲美甚至超越有监督方法
  • 生成的声图可作特征提升下游模型表现,适合部署于多场景系统

声学映射技术长期用于声源到达方向估计(DoAE)。传统波束成形方法虽具可解释性,但依赖迭代求解,计算量大且对声学环境敏感;而近年的有监督深度学习方法虽具备前馈速度和鲁棒性,却需大量标注数据且缺乏可解释性。两者均难以在多样声学设置与阵列配置下稳定泛化。本文提出潜空间声学映射(LAM)框架,一种融合传统方法可解释性与深度学习适应性、效率优势的自监督方法。LAM可生成高分辨率声场映射,适应不同声学条件,并在多种麦克风阵列上高效运行。我们在LOCATA和STARSS基准上评估其鲁棒性,结果表明其定位性能达到或优于现有有监督方法。此外,我们证明LAM生成的声图可作为有效特征,进一步提升有监督模型的定位精度,凸显其在构建自适应、高性能声源定位系统中的潜力。

原文摘要 · Abstract (English)

Acoustic mapping techniques have long been used in spatial audio processing for direction of arrival estimation (DoAE). Traditional beamforming methods for acoustic mapping, while interpretable, often rely on iterative solvers that can be computationally intensive and sensitive to acoustic variability. On the other hand, recent supervised deep learning approaches offer feedforward speed and robustness but require large labeled datasets and lack interpretability. Despite their strengths, both methods struggle to consistently generalize across diverse acoustic setups and array configurations, limiting their broader applicability. We introduce the Latent Acoustic Mapping (LAM) model, a self-supervised framework that bridges the interpretability of traditional methods with the adaptability and efficiency of deep learning methods. LAM generates high-resolution acoustic maps, adapts to varying acoustic conditions, and operates efficiently across different microphone arrays. We assess its robustness on DoAE using the LOCATA and STARSS benchmarks. LAM achieves comparable or superior localization performance to existing supervised methods. Additionally, we show that LAM's acoustic maps can serve as effective features for supervised models, further enhancing DoAE accuracy and underscoring its potential to advance adaptive, high-performance sound localization systems.

声源定位自监督学习声场映射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。