arXiv:2608.10680cs.CV2026-08

解决可见光与红外图像严重错位问题,实现端到端精准目标检测。

Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration

论文配图:Bridging Severe Cross-Modal Misalignment: End-to-End Visible-Infrared Object Detection via Explicit Feature-Domain Affine Registration
图 1 · 摘自论文原文
  • 提出显式特征域仿射配准,直接校正跨模态几何偏差。
  • 在新构建的DVMA数据集上达到69.7% mAP50,性能领先。
  • 适合处理无人机等场景下的可见光-红外检测任务。

可见光-红外目标检测依赖于互补的RGB与热成像线索,但跨模态空间错位常导致性能下降。现有方法多采用隐式特征适应处理弱错位,对大偏移几何差异仍不足。本文提出面向严重跨模态几何失配的端到端可见光-红外检测网络JFRDet,引入跨模态仿射对齐(CMAA)模块,显式估计图像级仿射变换以实现多层次特征对齐。光照变化直接影响RGB线索可靠性,因此设计光照引导的互补融合(IGCF)模块,动态调整模态可信度进行跨模态融合。进一步提出对齐质量一致性门控(AQCG)策略,根据对齐可靠性与梯度一致性调节检测监督,稳定联合优化过程。我们还构建了用于评估严重跨模态几何失配下检测性能的新基准DVMA。所提JFRDet在DVMA上取得69.7% mAP50,达到当前最先进水平。代码与数据集将开源于GitHub。

原文摘要 · Abstract (English)

Visible-infrared object detection relies on complementary RGB and thermal cues, but its performance is often degraded by cross-modal spatial misalignment. Most existing methods rely on implicit feature adaptation to handle weakly misaligned scenarios, while large-offset geometric discrepancies remain insufficiently addressed. In this paper, we propose a Joint Feature-domain Registration and Detection network (JFRDet), an end-to-end visible-infrared oriented object detector tailored for severely cross-modal geometric discrepancies. JFRDet introduces a Cross-Modal Affine Alignment (CMAA) module to estimate an image-level affine transformation for explicit multi-level feature alignment. Note that illumination changes directly affect the reliability of RGB cues, an Illumination-Guided Complementary Fusion (IGCF) module adaptively exploits modality reliability under varying illumination conditions for cross-modal fusion. Then, an Alignment Quality-Consistency Gating (AQCG) strategy stabilizes joint optimization by modulating detection supervision according to alignment reliability and gradient consistency. We further construct DroneVehicle Misaligned (DVMA), a benchmark for evaluating visible-infrared oriented object detection under severe cross-modal geometric misalignment. The proposed JFRDet achieves 69.7\% $\mathrm{mAP}_{50}$ on DVMA, which represents state-of-the-art (SOTA) performance. The code and dataset will be available on GitHub.

跨模态检测红外视觉特征对齐无人机感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。