用数学距离衡量视觉编码器匹配度,提升视觉语言模型选型效率
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance

- 引入格罗莫夫-沃瑟斯坦距离量化跨模态结构相似性
- 该指标与最终模型性能相关性超传统指标三倍以上
- 无需完整训练即可预测最佳视觉编码器,适合快速原型设计
视觉语言模型(VLM)通过集成视觉编码器扩展了传统大语言模型的视觉理解能力。尽管已有研究尝试多种视觉编码器与大语言模型的组合,但对何种视觉编码器更适合VLM对齐仍缺乏系统性理解。本文通过对19个来自不同来源的预训练视觉编码器进行系统实验,发现常见选择标准(如模型最大或零样本准确率最高)均无法有效识别最优模型,这些指标与VLM性能仅呈弱至中等相关性。我们进一步揭示:跨模态结构相似性在视觉编码器选择中起关键作用,采用格罗莫夫-沃瑟斯坦距离作为度量代理。理论分析表明,跨模态映射的学习能力可被该距离严格关联。60余次全量VLM训练实验证明,该仅需推理的指标显著优于其他策略,且与最终性能相关性更强,可在不进行完整训练的情况下高效预测最优模型。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have enhanced traditional LLMs with visual capabilities through the integration of vision encoders. While recent works have explored various combinations of vision encoders and LLMs, there still lacks a principled understanding of what makes a vision encoder suitable for VLM alignment. In this paper, we systematically investigate this question via comprehensive experiments on a curated collection of 19 pre-trained vision encoders from diverse sources. We first demonstrate that common practices, such as choosing encoders with the largest size or highest zero-shot accuracy, consistently fail to identify optimal models. In fact, these metrics show only weak to moderate correlation with VLM performance. This intriguing finding begs a fundamental question: What factors of vision-encoders matter in VLM? Through comprehensive analysis, we identify that the structural similarity across modalities plays a crucial but previously overlooked role in vision-encoder selection, which we measure using the Gromov-Wasserstein distance as a proxy. From a theoretical perspective, we show that the learnability of cross-modality mapping can be provably associated with the Gromov-Wasserstein distance. Empirical verification on 60+ full VLM training runs shows that our proposed inference-only metric performs significantly better than alternative model selection strategies and exhibits a much stronger correlation with final VLM performance, thereby enabling efficient and effective prediction of VLM performance before full training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。