arXiv:2511.15312cs.CV2025-11被引 3

融合雷达、视频与音频的多模态模型,精准识别无人机

A Multimodal Transformer Approach for UAV Detection and Aerial Object Recognition Using Radar, Audio, and Video Data

  • 用Transformer融合雷达、视觉、红外和音频数据
  • 测试集上准确率0.9812,对无人机识别精度极高
  • 实时推理达41.11帧/秒,适合部署于安防系统

无人飞行器(UAV)检测与空中目标识别对现代监控与安全至关重要,亟需克服单模态方法局限性的鲁棒系统。本研究设计并严格评估了一种新型多模态Transformer模型,整合雷达、可见光视频(RGB)、红外视频和音频数据流。该架构通过Transformer自注意力机制有效融合各模态特征,学习全面、互补且高度判别性的表示用于分类。在独立测试集上表现优异,宏平均指标达:准确率0.9812,召回率0.9873,精确率0.9787,F1分数0.9826,特异性0.9954。尤其在区分无人机与其他空中物体时表现出极高的精确率与召回率。计算分析显示其高效性:1.09 GFLOPs,122万参数,推理速度达41.11 FPS,证实其适用于实时应用。本研究验证了基于Transformer的多模态数据融合在空中目标分类中的有效性,达到当前最优性能,为复杂空域中无人机检测与监控提供了高精度、强鲁棒的解决方案。

原文摘要 · Abstract (English)

Unmanned aerial vehicle (UAV) detection and aerial object recognition are critical for modern surveillance and security, prompting a need for robust systems that overcome limitations of single-modality approaches. This research addresses these challenges by designing and rigorously evaluating a novel multimodal Transformer model that integrates diverse data streams: radar, visual band video (RGB), infrared (IR) video, and audio. The architecture effectively fuses distinct features from each modality, leveraging the Transformer's self-attention mechanisms to learn comprehensive, complementary, and highly discriminative representations for classification. The model demonstrated exceptional performance on an independent test set, achieving macro-averaged metrics of 0.9812 accuracy, 0.9873 recall, 0.9787 precision, 0.9826 F1-score, and 0.9954 specificity. Notably, it exhibited particularly high precision and recall in distinguishing drones from other aerial objects. Furthermore, computational analysis confirmed its efficiency, with 1.09 GFLOPs, 1.22 million parameters, and an inference speed of 41.11 FPS, highlighting its suitability for real-time applications. This study presents a significant advancement in aerial object classification, validating the efficacy of multimodal data fusion via a Transformer architecture for achieving state-of-the-art performance, thereby offering a highly accurate and resilient solution for UAV detection and monitoring in complex airspace.

无人机检测多模态融合Transformer实时识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。