arXiv:2608.04833cs.CV2026-08

用注册令牌实现可见光与红外目标检测的高效融合

RegisterBridgeMM: A Register-Centric Framework for RGB-Infrared Object Detection

  • 以预训练注册令牌为中介,分三阶段完成跨模态信息融合
  • 在四个数据集上均达最高mAP50-95,且冻结主干网络
  • 适合需要轻量高效融合的多模态检测场景

可见光-红外(RGB-IR)目标检测得益于互补的可见光与热成像线索,但在光照变化、天气波动和复杂场景下,有效融合仍具挑战。现有方法常在表达力强的像素级交互与轻量但受限的适应机制间权衡。我们观察到,预训练的注册令牌在配对的RGB-IR输入中同时包含模态共享与模态特异性信息,可作为跨模态通信的紧凑基础。基于此,提出RegisterBridgeMM,一种以注册令牌为中心的三阶段融合框架:聚合阶段保留预训练的单模态注册摘要;桥接阶段通过双向注册-像素读取与共识残差调节实现跨模态交互;投影阶段将注册摘要转化为空间自适应的像素特征校准。该路径避免密集的像素-像素跨模态交互,同时保留预训练像素表示。在所有四个基准测试(LLVIP、M3FD、DroneVehicle、FLIR-Aligned)上,即使冻结主干网络,RegisterBridgeMM也实现了最高的mAP50-95。

原文摘要 · Abstract (English)

RGB-infrared (RGB-IR) object detection benefits from complementary visible and thermal cues, but effective fusion remains challenging under illumination changes, weather variation, and cluttered scenes. Existing RGB-IR fusion methods often trade expressive patch-level interaction for lighter but more constrained adaptation mechanisms. We empirically observe that pretrained register tokens contain both modality-shared and modality-specific information on paired RGB-IR inputs, suggesting that they can serve as a compact substrate for cross-modal communication. Building on this observation, we propose RegisterBridgeMM, a register-mediated fusion framework organized as a three-stage register lifecycle. Aggregate preserves per-modality register summarization inherited from pretraining; Bridge performs bidirectional register-to-patch reading with consensus-residual regulation; and Project translates the resulting register summary into spatially adaptive calibration of patch features. This register pathway avoids dense patch-to-patch cross-modal interaction while preserving the pretrained patch representation. With both backbone streams frozen, RegisterBridgeMM achieves the highest mAP50-95 among the evaluated methods on all four benchmarks: LLVIP, M3FD, DroneVehicle, and FLIR-Aligned.

目标检测多模态融合注册令牌红外感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。