arXiv:2509.12710cs.CV2025-09被引 1

用语言引导的红外可见光图像融合新方法,提升语义对齐精度

RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation

  • 将文本驱动融合与指代图像分割结合,通过联合优化提升语义一致性
  • 在12.5k训练数据上实现超过11%的mIoU提升,达到当前最优
  • 提出新基准MM-RIS,适合多模态语义理解与图像融合研究者使用

文本驱动的红外与可见光图像融合因能通过自然语言指导融合过程而受到关注。然而,现有方法缺乏目标对齐的任务来监督和评估输入文本对融合结果的有效性。我们观察到,指代图像分割(RIS)与文本驱动融合具有共同目标:突出文本所指物体。受此启发,我们提出RIS-FUSION,一种级联框架,通过联合优化统一融合与RIS任务。核心是LangGatedFusion模块,将文本特征注入融合主干以增强语义对齐。为支持多模态指代图像分割任务,我们引入MM-RIS,一个包含12.5k训练和3.5k测试三元组的大规模基准,每个三元组包含一对红外-可见光图像、一个分割掩码和一个指代表达。大量实验表明,RIS-FUSION在mIoU上优于现有方法超过11%。代码与数据集将公开于https://github.com/SijuMa2003/RIS-FUSION。

原文摘要 · Abstract (English)

Text-driven infrared and visible image fusion has gained attention for enabling natural language to guide the fusion process. However, existing methods lack a goal-aligned task to supervise and evaluate how effectively the input text contributes to the fusion outcome. We observe that referring image segmentation (RIS) and text-driven fusion share a common objective: highlighting the object referred to by the text. Motivated by this, we propose RIS-FUSION, a cascaded framework that unifies fusion and RIS through joint optimization. At its core is the LangGatedFusion module, which injects textual features into the fusion backbone to enhance semantic alignment. To support multimodal referring image segmentation task, we introduce MM-RIS, a large-scale benchmark with 12.5k training and 3.5k testing triplets, each consisting of an infrared-visible image pair, a segmentation mask, and a referring expression. Extensive experiments show that RIS-FUSION achieves state-of-the-art performance, outperforming existing methods by over 11% in mIoU. Code and dataset will be released at https://github.com/SijuMa2003/RIS-FUSION.

图像融合多模态语义对齐指代分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。