arXiv:2607.19040cs.CV2026-07

通过优先图引导,提升红外弱小无人机检测精度。

Gaze-DETR: Top-Down Guidance Through Priority Maps for Infrared Weak-Small UAV Detection with DETR

论文配图:Gaze-DETR: Top-Down Guidance Through Priority Maps for Infrared Weak-Small UAV Detection with DETR
图 1 · 摘自论文原文
  • 引入生物启发的优先图,在定位前筛选关键区域。
  • 在TIR-UAV120-Gaze上达86.18 mAP₅₀,Anti-UAV410上达87.08 mAP₅₀。
  • 适用于弱小目标检测,尤其适合缺乏标注数据场景。

红外弱小目标检测(ISTD)因目标微小、对比度低,易被杂波、噪声或遮挡淹没而面临挑战。传统单帧与多帧检测器依赖边界框监督,仅提供最终位置信息,难以指导候选区域优先级或保留弱目标特征。任务驱动视觉搜索可提供此类引导:自上而下的目标与视觉证据共同生成空间优先图以排序候选位置。基于此,我们提出Gaze-DETR,一种在定位前学习内部优先图的生物启发检测器。首先,优先头从图像特征预测归一化优先图;其次,残差优先引导特征调制(RPFM)增强高优先级响应并保留多尺度特征;最后,优先引导锚点查询注入(PAQI)将高优先级位置转为解码器锚点查询。使用三种监督方案训练优先头:基于框的高斯图、由注视密度图构建的真实注视图,以及从配对标注中学习的转移伪注视图,并应用于Anti-UAV410训练框。为支持后两种方案,我们构建了含配对检测与任务驱动眼动标注的TIR-UAV120-Gaze数据集。在TIR-UAV120-Gaze上,使用真实注视监督达86.18 mAP₅₀与89.00 F1;在Anti-UAV410上,使用转移伪注视监督达87.08 mAP₅₀与90.43 F1。结果表明,显式空间优先学习在不同标注设置下均提供互补于边界框监督的预定位引导。

原文摘要 · Abstract (English)

Infrared small target detection (ISTD) remains challenging because tiny, low-contrast targets are easily overwhelmed by clutter, noise, or occlusion. Conventional single-frame and multi-frame detectors rely on bounding-box supervision, which specifies final target locations but offers little explicit guidance for prioritizing candidate regions or preserving weak-target evidence before localization. Task-driven visual search offers such guidance: top-down goals and visual evidence jointly form a spatial priority map that ranks candidate locations. Building on this principle, we propose Gaze-DETR, a bio-inspired detector that learns an internal priority map before localization. First, a priority head predicts a normalized priority map from image features. Second, Residual Priority-Guided Feature Modulation (RPFM) enhances high-priority responses while retaining multi-scale features. Finally, Priority-Guided Anchor Query Injection (PAQI) converts high-priority locations into decoder anchor queries. We train the priority head using three supervision schemes: box-derived Gaussian maps; real-gaze maps constructed from fixation-density maps; and transferred pseudo-gaze maps learned from gaze--box relations in paired annotations and applied to Anti-UAV410 training boxes. To support the latter two schemes, we construct TIR-UAV120-Gaze with paired detection and task-driven eye-tracking annotations. On TIR-UAV120-Gaze, Gaze-DETR achieves 85.76 mAP$_{50}$ and 88.77 F1 with box-derived supervision, and 86.18 mAP$_{50}$ and 89.00 F1 with real-gaze supervision. On Anti-UAV410, it achieves 87.06 mAP$_{50}$ and 90.90 F1 with box-derived supervision, and 87.08 mAP$_{50}$ and 90.43 F1 with transferred pseudo-gaze supervision. These results show that explicit spatial-priority learning provides pre-localization guidance complementary to bounding-box supervision across annotation settings and costs.

红外检测小目标注意力机制生物启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。