让视觉模型更懂语言,还能适应不同分辨率图像。
Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
- 用持续位置编码处理任意分辨率图像,同时对齐图文表征。
- 在DINOv2、SigLIP等模型上提升多模态理解能力,不损失原有性能。
- 适合想提升视觉模型跨模态能力的研究者和工程师。
视觉基础模型(VFMs)为众多应用提供了强大的视觉表征能力。本文通过多模态持续预训练增强现有VFMs,使其能够有效处理不同分辨率的视觉输入,并生成与语言表征更对齐的视觉表征,且不受原始预训练目标限制。为此,提出M-CPT框架:引入持续位置编码(CPE)以灵活处理视觉分辨率变化,设计特征对齐目标,在多模态训练中提升视觉与文本表征的一致性。在DINOv2、SigLIP和AIMv2等主流视觉基础模型上的大量实验表明,M-CPT在保持分类、分割等标准视觉基准性能的同时,持续提升了多模态理解表现。
原文摘要 · Abstract (English)
Vision Foundation Models (VFMs) provide strong visual representations for a wide range of applications. In this work, we enhance prevailing VFMs through multimodal training, allowing them to effectively process visual inputs at varying resolutions while producing visual representations that are better aligned with language representations, regardless of their original pre-training objectives. To this end, we introduce M-CPT, a Multimodal Continual Pre-Training framework designed to improve the understanding capability of pre-trained VFMs while preserving their strong visual representation quality. M-CPT introduces a Continual Position Embedding (CPE) for handling flexible visual resolutions, along with a feature alignment objective that improves the consistency between visual and textual representations during multimodal training. Extensive experiments on leading VFMs, including DINOv2, SigLIP, and AIMv2, demonstrate that M-CPT consistently improves multimodal understanding performance while preserving strong performance on standard vision benchmarks such as classification and segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。