arXiv:2511.06406cs.CVcs.AI2025-11被引 1

提出可适配单/双模态的红外可见光目标检测新架构

On Modality Incomplete Infrared-Visible Object Detection: An Architecture Compatibility Perspective

  • 设计无模态依赖的可插拔注意力模块,支持灵活模态切换
  • 在缺失主模态时仍保持高检测精度,优于现有方法
  • 适用于全天候应用,尤其适合传感器不全的场景

红外与可见光目标检测(IVOD)在全天候应用中至关重要。尽管进展显著,现有模型在面对不完整模态数据时性能明显下降,尤其是主导模态缺失时。本文从架构兼容性视角深入研究该问题,提出一种即插即用的Scarf Neck模块,专用于DETR系列模型,引入无模态依赖的可变形注意力机制,使检测器在训练和推理时能灵活适应单模态或双模态输入。训练时采用伪模态丢弃策略,充分挖掘多模态信息,提升对单/双模态工作模式的兼容性与鲁棒性。此外,构建了全面的基准测试集,系统评估缺失模态为主或为次的情况。所提Scarf-DETR不仅在缺失模态场景表现优异,还在标准完整模态基准上取得更优结果。代码将公开于https://github.com/YinghuiXing/Scarf-DETR。

原文摘要 · Abstract (English)

Infrared and visible object detection (IVOD) is essential for numerous around-the-clock applications. Despite notable advancements, current IVOD models exhibit notable performance declines when confronted with incomplete modality data, particularly if the dominant modality is missing. In this paper, we take a thorough investigation on modality incomplete IVOD problem from an architecture compatibility perspective. Specifically, we propose a plug-and-play Scarf Neck module for DETR variants, which introduces a modality-agnostic deformable attention mechanism to enable the IVOD detector to flexibly adapt to any single or double modalities during training and inference. When training Scarf-DETR, we design a pseudo modality dropout strategy to fully utilize the multi-modality information, making the detector compatible and robust to both working modes of single and double modalities. Moreover, we introduce a comprehensive benchmark for the modality-incomplete IVOD task aimed at thoroughly assessing situations where the absent modality is either dominant or secondary. Our proposed Scarf-DETR not only performs excellently in missing modality scenarios but also achieves superior performances on the standard IVOD modality complete benchmarks. Our code will be available at https://github.com/YinghuiXing/Scarf-DETR.

目标检测多模态DETR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。