发现视觉语言模型的语法理解能力弱于纯语言模型。
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
- 对比不同模型的语法编码能力,发现视觉语言模型表现更差。
- 预训练目标是影响语法学习的关键因素,超过模型大小和数据量。
- 中间层更擅长语法编码,而CLIP模型的语法能力随层数下降。
视觉语言模型(VLM)是图像描述、文生图等多模态应用的基础模型。尽管已有研究指出其文本编码器在组合性与语义理解方面存在局限,但根本原因仍不明确。本文通过分析VLM文本编码器对句法信息的编码能力,比较了不同目标函数、参数规模、训练数据量的VLM,以及单模态语言模型(ULM)的表现。结果表明,ULM的文本编码器比VLM更有效获取句法知识。VLM的句法学习主要受预训练目标影响,其作用超过模型架构、规模或训练数据量。不同模型呈现不同的层间趋势:CLIP的句法能力随层数递减,而其他模型在中间层表现出更强的句法编码能力。
原文摘要 · Abstract (English)
Vision-language models (VLMs), serve as foundation models for multi-modal applications such as image captioning and text-to-image generation. Recent studies have highlighted limitations in VLM text encoders, particularly in areas like compositionality and semantic understanding, though the underlying reasons for these limitations remain unclear. In this work, we aim to address this gap by analyzing the syntactic information, one of the fundamental linguistic properties, encoded by the text encoders of VLMs. We perform a thorough analysis comparing VLMs with different objective functions, parameter size and training data size, and with uni-modal language models (ULMs) in their ability to encode syntactic knowledge. Our findings suggest that ULM text encoders acquire syntactic information more effectively than those in VLMs. The syntactic information learned by VLM text encoders is shaped primarily by the pre-training objective, which plays a more crucial role than other factors such as model architecture, model size, or the volume of pre-training data. Models exhibit different layer-wise trends where CLIP performance dropped across layers while for other models, middle layers are rich in encoding syntactic knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。