arXiv:2412.12698cs.ROcs.SD2024-12中稿 · ICASSP被引 13

用音频和激光雷达伪标签,让系统仅凭声音就能准确定位无人机3D轨迹。

Audio Array-Based 3D UAV Trajectory Estimation with LiDAR Pseudo-Labeling

  • 音频转频谱图,用自监督学习训练模型,无需真实标注数据。
  • 激光雷达先估算轨迹作伪标签,指导音频模型学习,精度达新高。
  • 部署时只需音频输入,适合隐私敏感或无激光雷达场景。

随着小型无人机日益普及,其对公共安全与隐私的潜在影响引发关注,亟需先进的追踪与轨迹估计方案。本文提出一种基于音频阵列的三维无人机轨迹估计框架。首先将音频数据转换为梅尔频谱图,通过编码器提取关键时序与频谱信息;同时利用无监督方法从激光雷达点云中估计无人机轨迹,生成伪标签。该激光雷达系统作为教师网络,指导音频感知网络(学生网络)进行训练。训练完成后,模型仅需音频信号即可独立预测三维轨迹,无需激光雷达或外部真值。为进一步提升精度,引入高斯过程建模优化时空追踪性能。在MMAUD数据集上,本方法表现优异,成为无需真实标注的自监督学习轨迹估计新基准。

原文摘要 · Abstract (English)

As small unmanned aerial vehicles (UAVs) become increasingly prevalent, there is growing concern regarding their impact on public safety and privacy, highlighting the need for advanced tracking and trajectory estimation solutions. In response, this paper introduces a novel framework that utilizes audio array for 3D UAV trajectory estimation. Our approach incorporates a self-supervised learning model, starting with the conversion of audio data into mel-spectrograms, which are analyzed through an encoder to extract crucial temporal and spectral information. Simultaneously, UAV trajectories are estimated using LiDAR point clouds via unsupervised methods. These LiDAR-based estimations act as pseudo labels, enabling the training of an Audio Perception Network without requiring labeled data. In this architecture, the LiDAR-based system operates as the Teacher Network, guiding the Audio Perception Network, which serves as the Student Network. Once trained, the model can independently predict 3D trajectories using only audio signals, with no need for LiDAR data or external ground truth during deployment. To further enhance precision, we apply Gaussian Process modeling for improved spatiotemporal tracking. Our method delivers top-tier performance on the MMAUD dataset, establishing a new benchmark in trajectory estimation using self-supervised learning techniques without reliance on ground truth annotations.

无人机追踪音频感知自监督学习3D定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。