arXiv:2512.02668cs.CV2025-12被引 2

统一多模态反无人机跟踪框架,提升精度与效率

UAUTrack: Towards Unified Multimodal Anti-UAV Visual Tracking

  • 单流单阶段端到端架构融合多模态数据
  • 在Anti-UAV和DUT Anti-UAV上达最优性能
  • 文本提示引导模型专注识别无人机场景

反无人机(Anti-UAV)跟踪研究已探索多种模态,包括RGB、TIR及RGB-T融合。然而,跨模态协作的统一框架仍不完善。现有方法多采用独立模型处理单一任务,忽视跨模态信息共享潜力。此外,反无人机跟踪技术尚处初期,现有方案难以实现有效多模态融合。为此,本文提出UAUTrack,一种基于单流、单阶段、端到端架构的统一单目标跟踪框架,可高效整合多模态信息。该框架引入关键组件:文本先验提示策略,引导模型在多样化场景中聚焦无人机。实验结果表明,UAUTrack在Anti-UAV和DUT Anti-UAV数据集上达到当前最优性能,并在Anti-UAV410数据集上保持精度与速度的良好平衡,展现出在多样反无人机场景中的高精度与实用效率。

原文摘要 · Abstract (English)

Research in Anti-UAV (Unmanned Aerial Vehicle) tracking has explored various modalities, including RGB, TIR, and RGB-T fusion. However, a unified framework for cross-modal collaboration is still lacking. Existing approaches have primarily focused on independent models for individual tasks, often overlooking the potential for cross-modal information sharing. Furthermore, Anti-UAV tracking techniques are still in their infancy, with current solutions struggling to achieve effective multimodal data fusion. To address these challenges, we propose UAUTrack, a unified single-target tracking framework built upon a single-stream, single-stage, end-to-end architecture that effectively integrates multiple modalities. UAUTrack introduces a key component: a text prior prompt strategy that directs the model to focus on UAVs across various scenarios. Experimental results show that UAUTrack achieves state-of-the-art performance on the Anti-UAV and DUT Anti-UAV datasets, and maintains a favourable trade-off between accuracy and speed on the Anti-UAV410 dataset, demonstrating both high accuracy and practical efficiency across diverse Anti-UAV scenarios.

反无人机多模态视觉跟踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。