arXiv:2412.16928cs.SDcs.CV2024-12被引 12

用音视频融合实现轻量级无人机轨迹估计与分类

AV-DTEC: Self-Supervised Audio-Visual Fusion for Drone Trajectory Estimation and Classification

  • 自监督学习结合激光雷达生成标签,实现音视频特征并行提取
  • 在真实多模态数据中达到高精度,跨光照条件鲁棒性强
  • 适合需要低成本反无人机系统的安全防护场景

小型无人机的普及对公共安全构成严重威胁,而传统检测系统往往体积大且成本高。为此,本文提出 AV-DTEC,一种基于自监督音视频融合的轻量级反无人机系统。该系统利用激光雷达生成的标签进行自监督训练,通过并行选择性状态空间模型同时学习音频与视觉特征。为增强跨光照条件下的鲁棒性,设计了即插即用的主-辅特征增强模块,将视觉特征融入音频特征。为进一步减少对辅助特征依赖并对齐模态,提出教师-学生模型以自适应调整视觉特征权重。AV-DTEC 在真实多模态数据上表现出卓越的准确性和有效性。代码与训练模型已公开于 GitHub。

原文摘要 · Abstract (English)

The increasing use of compact UAVs has created significant threats to public safety, while traditional drone detection systems are often bulky and costly. To address these challenges, we propose AV-DTEC, a lightweight self-supervised audio-visual fusion-based anti-UAV system. AV-DTEC is trained using self-supervised learning with labels generated by LiDAR, and it simultaneously learns audio and visual features through a parallel selective state-space model. With the learned features, a specially designed plug-and-play primary-auxiliary feature enhancement module integrates visual features into audio features for better robustness in cross-lighting conditions. To reduce reliance on auxiliary features and align modalities, we propose a teacher-student model that adaptively adjusts the weighting of visual features. AV-DTEC demonstrates exceptional accuracy and effectiveness in real-world multi-modality data. The code and trained models are publicly accessible on GitHub \url{https://github.com/AmazingDay1/AV-DETC}.

反无人机音视频融合自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。