用文本引导跨模态对齐,解决可见光与红外图像检测差异问题
Bridging the RGB-IR Gap: Consensus and Discrepancy Modeling for Text-Guided Multispectral Detection
- 以文本为桥梁实现可见光与红外图像的语义对齐
- 融合一致信息和互补差异,提升检测精度
- 适合多模态目标检测与遥感视觉研究者
文本引导的多光谱目标检测利用文本语义指导可见光(RGB)与红外(IR)之间的跨模态交互,实现更鲁棒的感知。然而现有方法存在两大局限:(1) 文本仅作为辅助语义增强信号,未充分发挥其在弥合RGB与IR固有粒度差异中的引导作用;(2) 常规数据驱动的注意力融合倾向于强调稳定的一致性,忽视潜在有价值的跨模态差异。为此,我们提出一种带有双支持建模的语义桥接融合框架。具体地,将文本作为共享语义桥梁,在统一类别条件下对齐RGB与IR响应,同时将重新校准的热成像语义先验投影至RGB分支进行语义级映射融合。进一步将RGB-IR交互证据形式化为常规共识支持与包含潜在判别性线索的互补差异支持,并通过动态重校准引入融合,形成结构化的归纳偏置。此外,设计双向语义对齐模块实现闭环视觉-文本引导增强。大量实验验证了所提框架的有效性及其在多光谱基准上的优越检测性能。代码已开源。
原文摘要 · Abstract (English)
Text-guided multispectral object detection uses text semantics to guide semantic-aware cross-modal interaction between RGB and IR for more robust perception. However, notable limitations remain: (1) existing methods often use text only as an auxiliary semantic enhancement signal, without exploiting its guiding role to bridge the inherent granularity asymmetry between RGB and IR; and (2) conventional data-driven attention-based fusion tends to emphasize stable consensus while overlooking potentially valuable cross-modal discrepancies. To address these issues, we propose a semantic bridge fusion framework with bi-support modeling for multispectral object detection. Specifically, text is used as a shared semantic bridge to align RGB and IR responses under a unified category condition, while the recalibrated thermal semantic prior is projected onto the RGB branch for semantic-level mapping fusion. We further formulate RGB-IR interaction evidence into the regular consensus support and the complementary discrepancy support that contains potentially discriminative cues, and introduce them into fusion via dynamic recalibration as a structured inductive bias. In addition, we design a bidirectional semantic alignment module for closed-loop vision-text guidance enhancement. Extensive experiments demonstrate the effectiveness of the proposed fusion framework and its superior detection performance on multispectral benchmarks. Code is available at https://github.com/zhenwang5372/Bridging-RGB-IR-Gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。