arXiv:2512.04581cs.CV2025-12中稿 · IEEE TMM

针对红外无人机追踪的弱特征难题,提出动态特征融合与注意力蒸馏新方法。

Infrared UAV Target Tracking with Dynamic Feature Refinement and Global Contextual Attention Knowledge Distillation

  • 设计动态特征融合网络,自适应增强目标区域特征。
  • 通过上下文注意力蒸馏,提升模型对关键区域的关注度。
  • 在复杂背景下实现高精度实时追踪,适合安防反制场景。

基于热成像的无人机红外目标追踪是反无人机应用中的关键技术。然而,红外无人机目标常呈现特征微弱、背景复杂等问题,给精确追踪带来挑战。为此,本文提出SiamDFF——一种融合特征增强与全局上下文注意力知识蒸馏的动态特征融合孪生网络。该方法包含选择性目标增强网络(STEN)、动态空间特征聚合模块(DSFAM)和动态通道特征聚合模块(DCFAM)。STEN利用强度感知多头交叉注意力,自适应增强模板与搜索分支的重要区域;DSFAM通过空间注意力引导,将局部细节与全局特征融合,强化多尺度目标特征;DCFAM整合经STEN生成的混合模板与原始模板,减少背景干扰,突出目标区域特征。此外,提出一种新型面向追踪任务的目标感知上下文注意力知识蒸馏器,将教师网络中的目标先验传递至学生模型,在不增加计算开销的前提下,显著提升学生网络在骨干网络各层级对信息区域的关注能力。在真实红外无人机数据集上的大量实验表明,该方法在复杂背景下优于现有最先进追踪算法,同时保持实时追踪速度。

原文摘要 · Abstract (English)

Unmanned aerial vehicle (UAV) target tracking based on thermal infrared imaging has been one of the most important sensing technologies in anti-UAV applications. However, the infrared UAV targets often exhibit weak features and complex backgrounds, posing significant challenges to accurate tracking. To address these problems, we introduce SiamDFF, a novel dynamic feature fusion Siamese network that integrates feature enhancement and global contextual attention knowledge distillation for infrared UAV target (IRUT) tracking. The SiamDFF incorporates a selective target enhancement network (STEN), a dynamic spatial feature aggregation module (DSFAM), and a dynamic channel feature aggregation module (DCFAM). The STEN employs intensity-aware multi-head cross-attention to adaptively enhance important regions for both template and search branches. The DSFAM enhances multi-scale UAV target features by integrating local details with global features, utilizing spatial attention guidance within the search frame. The DCFAM effectively integrates the mixed template generated from STEN in the template branch and original template, avoiding excessive background interference with the template and thereby enhancing the emphasis on UAV target region features within the search frame. Furthermore, to enhance the feature extraction capabilities of the network for IRUT without adding extra computational burden, we propose a novel tracking-specific target-aware contextual attention knowledge distiller. It transfers the target prior from the teacher network to the student model, significantly improving the student network's focus on informative regions at each hierarchical level of the backbone network. Extensive experiments on real infrared UAV datasets demonstrate that the proposed approach outperforms state-of-the-art target trackers under complex backgrounds while achieving a real-time tracking speed.

红外追踪孪生网络知识蒸馏无人机反制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。