对比学习视觉语言模型在长句理解与组合推理间存在敏感依赖关系。
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
- 通过控制实验分析两种能力的迁移条件
- 高质量长句数据能同步提升两种能力
- 冻结位置编码可能抑制组合推理
对比学习视觉语言模型(VLMs)在关联视觉与文本信息方面取得显著进展,但对长篇、组合性描述的理解仍是开放挑战。尽管这两项能力常被假设密切相关,其相互促进的具体条件仍不明确。本文通过跨多种训练目标、数据集和架构设计的受控实验,系统分析组合推理与长句理解之间的双向迁移机制。结果表明二者存在敏感的双向关系:在低质量标注或参数更新受限的训练条件下,模型无法泛化;而使用强视觉对齐的高质量长句数据则能同时促进两项能力。此外,为保持通用对齐而采用的冻结位置编码等架构设计,反而可能抑制组合学习。研究为数据选择与模型设计提供了可操作的优化建议。
原文摘要 · Abstract (English)
Contrastive vision-language models (VLMs) have made significant progress in binding visual and textual information, yet understanding long, compositional captions remains an open challenge. While these capabilities are often assumed to be closely related, the conditions under which they reinforce each other remain unclear. In this paper, we empirically analyze when compositional reasoning and long-caption understanding transfer across tasks, and when this relationship fails. Through controlled experiments across diverse training objectives, datasets, and architectural designs, we find a bidirectional but sensitive relationship between the two capabilities. Models trained on poorly grounded captions or with limited parameter updates fail to generalize, while high-quality long-caption data with strong visual grounding promotes both capabilities simultaneously. We further show that architectural choices aimed at preserving general alignment, such as frozen positional embeddings, can inadvertently limit compositional learning. Our analysis provides actionable guidelines for data selection and model design to improve VLM generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。