用大模型先验让红外可见图像融合更适配下游视觉任务
MambaTrans: Multimodal Fusion Image Translation via Large Language Model Priors for Downstream Visual Tasks
- 用大模型描述和分割掩码引导图像模态转换
- 在不修改预训练模型前提下提升检测与分割性能
- 适合需要融合多模态图像的智能感知系统
多模态图像融合旨在整合红外与可见光图像的互补信息,生成适用于下游任务的融合图像。现有下游预训练模型通常基于可见光图像训练,但可见光与多模态融合图像间存在显著像素分布差异,导致下游任务性能下降,甚至低于仅使用可见光图像的结果。本文研究如何将具有显著模态差异的多模态融合图像适配至基于可见光训练的目标检测与语义分割模型。为此,提出MambaTrans,一种新型多模态融合图像模态转换器。MambaTrans以多模态大语言模型的描述和语义分割模型的掩码为输入,核心组件为多模态状态空间块,结合掩码-图像-文本交叉注意力与3D选择性扫描模块,增强纯视觉能力。通过利用目标检测先验知识,在训练中最小化检测损失,并捕捉文本、掩码与图像间的长期依赖关系。该方法使预训练模型无需参数调整即可获得良好效果。公开数据集上的实验表明,MambaTrans能有效提升多模态图像在下游任务中的表现。
原文摘要 · Abstract (English)
The goal of multimodal image fusion is to integrate complementary information from infrared and visible images, generating multimodal fused images for downstream tasks. Existing downstream pre-training models are typically trained on visible images. However, the significant pixel distribution differences between visible and multimodal fusion images can degrade downstream task performance, sometimes even below that of using only visible images. This paper explores adapting multimodal fused images with significant modality differences to object detection and semantic segmentation models trained on visible images. To address this, we propose MambaTrans, a novel multimodal fusion image modality translator. MambaTrans uses descriptions from a multimodal large language model and masks from semantic segmentation models as input. Its core component, the Multi-Model State Space Block, combines mask-image-text cross-attention and a 3D-Selective Scan Module, enhancing pure visual capabilities. By leveraging object detection prior knowledge, MambaTrans minimizes detection loss during training and captures long-term dependencies among text, masks, and images. This enables favorable results in pre-trained models without adjusting their parameters. Experiments on public datasets show that MambaTrans effectively improves multimodal image performance in downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。