arXiv:2409.00973cs.CV2024-09被引 2

提出通用红外可见光融合框架,跨任务性能领先。

IVGF: The Fusion-Guided Infrared and Visible General Framework

  • 基于SOTA基础模型,设计特征与令牌增强模块
  • 在两个数据集上均超越现有双模态方法
  • 适合需要多任务融合的遥感/安防场景

红外与可见光双模态任务如语义分割和目标检测,可通过融合互补信息在极端场景下实现鲁棒性能。当前多数方法采用特定任务框架,跨任务泛化能力受限。本文提出融合引导的红外-可见光通用框架IVGF,可轻松扩展至多种高层视觉任务。首先,采用当前最优的红外与可见光基础模型提取通用表征;其次,为丰富高层视觉任务的语义信息,分别设计特征增强模块与令牌增强模块;此外,提出注意力引导融合模块,有效挖掘两模态的互补特性;同时采用Cutout&Mix数据增强策略,进一步提升模型对区域互补性的挖掘能力。大量实验表明,IVGF在语义分割与目标检测任务中均优于现有先进双模态方法。详细消融实验证明各模块有效性,另一项实验还评估了该方法在双模态语义分割中对抗模态缺失的能力。

原文摘要 · Abstract (English)

Infrared and visible dual-modality tasks such as semantic segmentation and object detection can achieve robust performance even in extreme scenes by fusing complementary information. Most current methods design task-specific frameworks, which are limited in generalization across multiple tasks. In this paper, we propose a fusion-guided infrared and visible general framework, IVGF, which can be easily extended to many high-level vision tasks. Firstly, we adopt the SOTA infrared and visible foundation models to extract the general representations. Then, to enrich the semantics information of these general representations for high-level vision tasks, we design the feature enhancement module and token enhancement module for feature maps and tokens, respectively. Besides, the attention-guided fusion module is proposed for effectively fusing by exploring the complementary information of two modalities. Moreover, we also adopt the cutout&mix augmentation strategy to conduct the data augmentation, which further improves the ability of the model to mine the regional complementary between the two modalities. Extensive experiments show that the IVGF outperforms state-of-the-art dual-modality methods in the semantic segmentation and object detection tasks. The detailed ablation studies demonstrate the effectiveness of each module, and another experiment explores the anti-missing modality ability of the proposed method in the dual-modality semantic segmentation task.

多模态融合红外可见光通用框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。