UNIV模型提升红外可见光跨模态感知,解决传感器模式偏差问题。
UNIV: Unified Foundation Model for Infrared and Visible Modalities
- 采用自监督对比学习构建统一跨模态特征空间。
- 红外语义分割提升1.7 mIoU,目标检测提升0.7 mAP。
- 适用于复杂光照与天气下的多模态视觉任务研究者。
RGB与红外联合感知对于在多样气象和光照条件下实现鲁棒性至关重要。尽管基础模型在单一模态上表现优异,但面临显著的跨模态性能退化,我们将其归因于一种模式捷径——即优先关注表层传感器模式而非深层语义。为此,我们提出UNIV,一个面向红外与可见光模态的统一基础模型。其核心是补丁级跨模态对比学习(PCCL),一种自监督对比学习策略,用于构建统一的跨模态特征空间。PCCL利用冻结的预训练模型基于语义相似性采样伪补丁对,并通过吸引语义相关对、排斥无关对来对齐红外-可见表示。该过程同时增强跨模态对齐与类间语义可分性,引导模型聚焦于语义结构而非陷入模式捷径。为进一步支持跨模态学习,我们构建了当前最全面的可见光-红外基准数据集MVIP,包含98,992对精确配准图像,覆盖多样化场景。大量实验证明,UNIV在红外任务上表现卓越(语义分割提升1.7 mIoU,检测提升0.7 mAP),同时保持对RGB任务的竞争力。
原文摘要 · Abstract (English)
Joint RGB-infrared perception is essential for achieving robustness under diverse weather and illumination conditions. Although foundation models excel within single modalities, they suffer from substantial cross-modal degradation, an issue we attribute to a pattern shortcut, i.e., a modal bias that prioritizes superficial sensor patterns over underlying semantics. To address this problem, we introduce UNIV, a Unified foundation model for Infrared and Visible modalities. At the core of UNIV lies Patch Cross-modal Contrastive Learning (PCCL), a self-supervised contrastive learning strategy that constructs a unified cross-modal feature space. PCCL employs a frozen pre-trained model to sample pseudo patch pairs based on semantic similarity, and aligns infrared-visible representations by attracting semantically related pairs while repelling unrelated ones. This process simultaneously enhances cross-modal alignment and inter-class semantic separability, guiding the model to focus on semantic structure rather than falling into pattern shortcuts. To further enable cross-modal learning, we introduce MVIP, the most comprehensive visible-infrared benchmark to date, containing 98,992 precisely aligned image pairs across diverse scenes. Extensive experiments demonstrate UNIV's superior performance on infrared tasks (+1.7 mIoU for semantic segmentation and +0.7 mAP for detection), while maintaining competitive accuracy on RGB tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。