用原型图融合提升红外可见光目标检测,更准更聚焦。
ProtoHGF-Net: Prototype HyperGraph Fusion with Intra-modal Calibration for RGBT Object Detection

- 用原型级语义空间替代密集跨模态交互,聚焦目标特征。
- 在三个数据集上分别达到85.9%、88.2%、79.1%的mAP50,性能领先。
- 适合做复杂场景下多模态目标检测的研究与应用。
RGB-热成像(RGBT)目标检测通过融合可见纹理与热线索,在复杂场景中实现鲁棒感知。然而,现有方法主要依赖全分辨率特征上的密集跨模态交互,不可避免引入背景干扰,阻碍目标相关表征的学习。本文提出原型超图融合网络(ProtoHGF-Net),将跨模态融合重新定义为原型级语义交互,而非密集交互。具体地,设计原型超图融合机制,在紧凑的原型级语义空间中执行跨模态交互,实现对目标相关原型的更优选择性融合。为支持该机制,提出教师-掩码校准蒸馏策略,利用模态特异性教师和目标感知掩码,在融合前校准模态特征,抑制背景主导响应,生成更聚焦目标的特征。在DroneVehicle、DVTOD和FLIR三个数据集上的实验表明,ProtoHGF-Net分别取得85.9%、88.2%、79.1%的mAP50,达到当前最优性能。代码已开源。
原文摘要 · Abstract (English)
RGB-Thermal (RGBT) object detection enables robust perception in complex scenes by leveraging the complementary strengths of visible textures and thermal cues. However, existing methods mainly rely on dense cross-modal interactions over full-resolution features, which inevitably introduce background interference and hinder the learning of target-relevant representations. In this paper, we propose the Prototype HyperGraph Fusion Network (ProtoHGF-Net), a novel framework that redefines cross-modal fusion as prototype-level semantic interaction rather than the dense cross-modal interaction paradigm. Specifically, we design Prototype HyperGraph Fusion to perform cross-modal interaction in a compact prototype-level semantic space. This design enables more selective fusion among target-relevant prototypes. To support this prototype-level fusion, we propose Teacher-Mask Calibration Distillation, which calibrates modality features before fusion using modality-specific teachers and target-aware masks. This strategy suppresses backgrou- nd-dominant responses and produces more target-focused features. Extensive experiments on DroneVehicle, DVTOD, and FLIR demonstrate that ProtoHGF-Net achieves state-of-the-art performance with 85.9\% $mAP_{50}$, 88.2\% $mAP_{50}$, and 79.1\% $mAP_{50}$, respectively. Our code is available at \href{https://github.com/ZiMo-Chen/ProtoHGF}{GitHub}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。