arXiv:2411.13179cs.SDcs.CV2024-11被引 2

用仿真音频训练神经网络,显著提升实时声源定位精度。

SONNET: Enhancing Time Delay Estimation by Leveraging Simulated Audio

  • 基于仿真数据训练神经网络,避免真实标注数据不足
  • 在新场景真实数据上性能超越传统方法,无需重新训练
  • 可直接用于声源定位、自校准等下游任务

时间差估计(时延估计或到达时间差)是多点定位、波达方向估计和自校准等应用的关键。该任务旨在估计信号到达两个不同传感器的时间差。对于音频传感器,当前多数系统依赖经典方法如广义互相关相位变换(GCC-PHAT)。本文证明,即使仅使用合成数据训练,基于学习的方法也能在新真实数据上显著优于GCC-PHAT。为克服真实数据缺乏标注的问题,我们利用一个规模大且多样化的仿真数据集进行训练,该数据集充分捕捉了真实世界问题的相关特征。我们提供了可实时运行的模型SONNET(仿真优化的神经网络时移估计器),能直接应用于多种真实场景,无需再训练。实验表明,使用该模型在自校准等下游任务中表现远超传统方法。

原文摘要 · Abstract (English)

Time delay estimation or Time-Difference-Of-Arrival estimates is a critical component for multiple localization applications such as multilateration, direction of arrival, and self-calibration. The task is to estimate the time difference between a signal arriving at two different sensors. For the audio sensor modality, most current systems are based on classical methods such as the Generalized Cross-Correlation Phase Transform (GCC-PHAT) method. In this paper we demonstrate that learning based methods can, even based on synthetic data, significantly outperform GCC-PHAT on novel real world data. To overcome the lack of data with ground truth for the task, we train our model on a simulated dataset which is sufficiently large and varied, and that captures the relevant characteristics of the real world problem. We provide our trained model, SONNET (Simulation Optimized Neural Network Estimator of Timeshifts), which is runnable in real-time and works on novel data out of the box for many real data applications, i.e. without re-training. We further demonstrate greatly improved performance on the downstream task of self-calibration when using our model compared to classical methods.

声源定位神经网络仿真训练时延估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。