融合激光与毫米波雷达数据,提升低空无人机轨迹预测精度
Efficient UAV trajectory prediction: A multi-modal deep diffusion framework
- 用双模态特征提取+双向交叉注意力融合激光与雷达点云数据
- 在MMAUD数据集上比基线模型轨迹预测准确率提升40%
- 适合关注低空安防与多传感器融合的工程师和研究者
为满足低空经济中对非法无人机管理的需求,提出一种基于激光雷达(LiDAR)与毫米波雷达信息融合的多模态无人机轨迹预测方法。设计了名为多模态深度融合框架(Multi-Modal Deep Fusion Framework)的深度融合网络,整体架构包含两个模态专用特征提取网络及双向交叉注意力融合模块,旨在充分挖掘LiDAR与雷达点云在空间几何结构与动态反射特性上的互补信息。特征提取阶段,采用独立但结构相同的特征编码器分别处理LiDAR与雷达数据。特征提取后进入双向交叉注意力机制阶段,实现两模态间的信息互补与语义对齐。为验证模型有效性,采用CVPR 2024 UG2+无人机追踪与位姿估计挑战赛使用的MMAUD数据集作为训练与测试数据集。实验结果表明,所提多模态融合模型显著提升轨迹预测精度,在该数据集上相较基线模型提升40%。此外,通过消融实验验证了不同损失函数与后处理策略对模型性能的提升作用。该模型能有效利用多模态数据,为低空经济中的非法无人机轨迹预测提供高效解决方案。
原文摘要 · Abstract (English)
To meet the requirements for managing unauthorized UAVs in the low-altitude economy, a multi-modal UAV trajectory prediction method based on the fusion of LiDAR and millimeter-wave radar information is proposed. A deep fusion network for multi-modal UAV trajectory prediction, termed the Multi-Modal Deep Fusion Framework, is designed. The overall architecture consists of two modality-specific feature extraction networks and a bidirectional cross-attention fusion module, aiming to fully exploit the complementary information of LiDAR and radar point clouds in spatial geometric structure and dynamic reflection characteristics. In the feature extraction stage, the model employs independent but structurally identical feature encoders for LiDAR and radar. After feature extraction, the model enters the Bidirectional Cross-Attention Mechanism stage to achieve information complementarity and semantic alignment between the two modalities. To verify the effectiveness of the proposed model, the MMAUD dataset used in the CVPR 2024 UG2+ UAV Tracking and Pose-Estimation Challenge is adopted as the training and testing dataset. Experimental results show that the proposed multi-modal fusion model significantly improves trajectory prediction accuracy, achieving a 40% improvement compared to the baseline model. In addition, ablation experiments are conducted to demonstrate the effectiveness of different loss functions and post-processing strategies in improving model performance. The proposed model can effectively utilize multi-modal data and provides an efficient solution for unauthorized UAV trajectory prediction in the low-altitude economy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。