只需单模态标注即可实现红外可见光目标检测,大幅降低标注成本。
DOD-SA: Infrared-Visible Decoupled Object Detection with Single-Modality Annotations
- 构建双分支网络,实现跨模态知识迁移与监督。
- 分三阶段训练提升伪标签质量,显著改善检测精度。
- 适合大规模遥感场景,尤其标注资源有限的项目。
红外-可见光目标检测在实际应用中展现出巨大潜力,可通过融合红外与可见光图像的互补信息实现全天候鲁棒感知。然而,现有方法通常需要双模态标注才能输出解耦检测结果,导致标注成本高,在大规模遥感应用中难以扩展。为此,我们提出一种基于单模态标注的红外-可见光解耦目标检测框架DOD-SA。其核心为单/双模态协同师生网络(CoSD-TSNet),包含单模态分支(SM-Branch)与双模态解耦分支(DMD-Branch),支持从有标注模态向无标注模态的知识迁移,并实现跨分支有效监督。为提升伪标签质量,引入渐进式自调训练策略(PaST),分三阶段训练:1)SM-Branch自训练;2)SM-Branch引导DMD-Branch学习;3)DMD-Branch精修。此外,设计伪标签分配器(PLA)以解决训练中的模态错配问题。
原文摘要 · Abstract (English)
Infrared-visible object detection has shown great potential in real-world applications, enabling robust all-day perception by leveraging the complementary information of infrared and visible images. However, existing methods typically require dual-modality annotations to output decoupled detection results, leading to high annotation costs and limiting scalability in large-scale remote sensing applications. To address this challenge, we propose a novel infrared-visible \textbf{D}ecoupled \textbf{O}bject \textbf{D}etection framework with \textbf{S}ingle-modality \textbf{A}nnotations, called DOD-SA. It is built upon a Single- and Dual-Modality Collaborative Teacher-Student Network (CoSD-TSNet), which consists of a single-modality branch (SM-Branch) and a dual-modality decoupled branch (DMD-Branch). This design enables cross-modality knowledge transfer from the labeled modality to the unlabeled modality, and facilitates effective cross-branch supervision. To further improve the quality of pseudo-labels, we introduce a Progressive and Self-Tuning Training Strategy (PaST) that trains the model in three stages: 1) SM-Branch self-training, 2) SM-Branch guiding the learning of DMD-Branch, and 3) DMD-Branch refinement. In addition, we design a Pseudo Label Assigner (PLA) to match labels across modalities, explicitly addressing modality misalignment during training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。