arXiv:2507.16318cs.CV2025-07ICCV被引 14

首个通用多模态视觉基础模型,实现红外可见光融合的自监督学习。

M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision

  • 提出跨模态结构稀疏性度量,指导渐进式掩码训练。
  • 在11个数据集上验证,4类任务泛化性能超越传统方法。
  • 适合多模态感知、自动驾驶等需要鲁棒环境理解的场景。

RGB-Thermal(RGBT)多光谱视觉对复杂环境中的鲁棒感知至关重要。现有方法多采用针对特定任务的手动定制模型,受人工先验、模态偏差和数据瓶颈限制。为此,我们首次构建了通用的RGBT多光谱基础模型M-SpecGene,旨在通过自监督方式从大规模广泛数据中学习模态不变表示。该模型为多光谱融合提供了新思路,并将以往的单任务研究整合到统一范式中。针对RGBT数据中信息分布不均的特点,提出跨模态结构稀疏性(CMSS)度量以量化两模态的信息密度,并设计GMM-CMSS渐进式掩码策略,实现由易到难、以物体为中心的预训练过程。大量实验验证了M-SpecGene在11个数据集上对4类下游任务的泛化能力。代码将开源于https://github.com/CalayZhou/M-SpecGene。

原文摘要 · Abstract (English)

RGB-Thermal (RGBT) multispectral vision is essential for robust perception in complex environments. Most RGBT tasks follow a case-by-case research paradigm, relying on manually customized models to learn task-oriented representations. Nevertheless, this paradigm is inherently constrained by artificial inductive bias, modality bias, and data bottleneck. To address these limitations, we make the initial attempt to build a Generalized RGBT MultiSpectral foundation model (M-SpecGene), which aims to learn modality-invariant representations from large-scale broad data in a self-supervised manner. M-SpecGene provides new insights into multispectral fusion and integrates prior case-by-case studies into a unified paradigm. Considering the unique characteristic of information imbalance in RGBT data, we introduce the Cross-Modality Structural Sparsity (CMSS) metric to quantify the information density across two modalities. Then we develop the GMM-CMSS progressive masking strategy to facilitate a flexible, easy-to-hard, and object-centric pre-training process. Comprehensive experiments validate M-SpecGene's generalizability across eleven datasets for four RGBT downstream tasks. The code will be available at https://github.com/CalayZhou/M-SpecGene.

多光谱基础模型自监督红外可见光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。