arXiv:2409.08083cs.CV2024-09

将视觉大模型迁移至非自然图像模态,显著提升分割性能。

SimMAT: Exploring Transferability from Vision Foundation Models to Any Image Modality

  • 设计通用迁移层MAT,实现自然图像大模型向多模态数据的跨域迁移
  • 在新构建的基准上平均提升分割mIoU至53.88%,较基线提升超30个百分点
  • 适用于偏振、红外等特殊成像领域,推动多模态感知技术落地

如ChatGPT和Sora等大规模训练的奠基模型已带来革命性社会影响。然而,许多领域的传感器难以收集与自然图像相当规模的数据来训练强健的奠基模型。为此,本文提出一种简单有效的框架SimMAT,以研究一个开放问题:从自然RGB图像训练的视觉奠基模型向具有不同物理特性的其他图像模态(如偏振)的可迁移性。SimMAT由模态无关的迁移层(MAT)和预训练奠基模型构成。我们将SimMAT应用于代表性视觉奠基模型Segment Anything Model(SAM),以支持任意评估的新图像模态。由于缺乏相关基准,我们构建了一个新基准来评估迁移学习性能。实验结果证实了将视觉奠基模型迁移到其他传感器的巨大潜力。具体而言,SimMAT可在所评估模态上将分割性能(mIoU)从22.15%平均提升至53.88%,并持续优于其他基线方法。我们希望SimMAT能提高对跨模态迁移学习的关注,助力各领域利用视觉奠基模型获得更优成果。

原文摘要 · Abstract (English)

Foundation models like ChatGPT and Sora that are trained on a huge scale of data have made a revolutionary social impact. However, it is extremely challenging for sensors in many different fields to collect similar scales of natural images to train strong foundation models. To this end, this work presents a simple and effective framework SimMAT to study an open problem: the transferability from vision foundation models trained on natural RGB images to other image modalities of different physical properties (e.g., polarization). SimMAT consists of a modality-agnostic transfer layer (MAT) and a pretrained foundation model. We apply SimMAT to a representative vision foundation model Segment Anything Model (SAM) to support any evaluated new image modality. Given the absence of relevant benchmarks, we construct a new benchmark to evaluate the transfer learning performance. Our experiments confirm the intriguing potential of transferring vision foundation models in enhancing other sensors' performance. Specifically, SimMAT can improve the segmentation performance (mIoU) from 22.15% to 53.88% on average for evaluated modalities and consistently outperforms other baselines. We hope that SimMAT can raise awareness of cross-modal transfer learning and benefit various fields for better results with vision foundation models.

跨模态迁移视觉大模型图像分割多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。