arXiv:2602.08484eess.AS2026-02

无需标注数据,用物理规律引导模型追踪声源方向。

Physics-Guided Variational Model for Unsupervised Sound Source Tracking

  • 用变分编码器+物理解码器,将几何约束注入隐空间
  • 真实数据实验中性能超越传统方法,接近有监督模型
  • 对麦克风阵列布局变化和位置误差鲁棒,适合实际场景

声源追踪通常依赖经典阵列处理算法,而机器学习方法多需精确的源位置标签,获取成本高或不现实。本文提出一种物理引导的变分模型,实现完全无监督的单声源追踪。该方法结合变分编码器与基于物理的解码器,通过解析推导的双麦克风时延似然,将几何约束注入隐空间。无需真值标签,模型可直接从麦克风阵列信号中学习估计声源方向。在真实数据上的实验表明,该方法优于传统基线,在准确率和计算复杂度上达到当前最优有监督模型水平。进一步验证其对不匹配阵列几何具有良好泛化能力,且对麦克风位置信息损坏具有强鲁棒性。最后,我们提出了扩展至多声源追踪的自然路径,并给出理论修改方案。

原文摘要 · Abstract (English)

Sound source tracking is commonly performed using classical array-processing algorithms, while machine-learning approaches typically rely on precise source position labels that are expensive or impractical to obtain. This paper introduces a physics-guided variational model capable of fully unsupervised single-source sound source tracking. The method combines a variational encoder with a physics-based decoder that injects geometric constraints into the latent space through analytically derived pairwise time-delay likelihoods. Without requiring ground-truth labels, the model learns to estimate source directions directly from microphone array signals. Experiments on real-world data demonstrate that the proposed approach outperforms traditional baselines and achieves accuracy and computational complexity comparable to state-of-the-art supervised models. We further show that the method generalizes well to mismatched array geometries and exhibits strong robustness to corrupted microphone position metadata. Finally, we outline a natural extension of the approach to multi-source tracking and present the theoretical modifications required to support it.

声源追踪无监督学习物理模型麦克风阵列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。