提出新基准评估视觉语言模型的组合能力,发现GPT-4o表现不如开源模型。
MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
- 构建人工标注的多模态组合性评测集,覆盖物体交互与复杂组合
- 发现主流模型在细粒度组合感知上存在缺陷,GPT-4o表现低于最佳开源模型
- 适用于评估和改进多模态模型的推理与组合能力研究
大型视觉语言模型(VLMs)显著提升了多模态理解能力,广泛应用于图像视频描述、视觉问答和跨模态检索等任务。然而,对模型组合性——即理解并生成已知视觉与文本成分新组合的能力——仍缺乏全面认识。现有基准仅粗略评估对象、关系和属性层面的组合性,忽视对物体交互、计数及复杂组合的深层推理。为此,我们提出MMCOMPOSITION,一个全新的人工标注基准,用于全面准确评估VLMs的组合性。通过该基准,我们量化分析主流VLMs的表现,意外发现GPT-4o的组合性优于部分开源模型,但整体仍存在局限。实验揭示了模型在细粒度组合感知与推理上的不足,为未来VLM设计与训练提供了改进方向。资源详见:https://hanghuacs.github.io/MMComposition/
原文摘要 · Abstract (English)
The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal understanding, enabling more sophisticated and accurate integration of visual and textual information across various tasks, including image and video captioning, visual question answering, and cross-modal retrieval. Despite VLMs' superior capabilities, researchers lack a comprehensive understanding of their compositionality -- the ability to understand and produce novel combinations of known visual and textual components. Prior benchmarks provide only a relatively rough compositionality evaluation from the perspectives of objects, relations, and attributes while neglecting deeper reasoning about object interactions, counting, and complex compositions. However, compositionality is a critical ability that facilitates coherent reasoning and understanding across modalities for VLMs. To address this limitation, we propose MMCOMPOSITION, a novel human-annotated benchmark for comprehensively and accurately evaluating VLMs' compositionality. Our proposed benchmark serves as a complement to these earlier works. With MMCOMPOSITION, we can quantify and explore the compositionality of the mainstream VLMs. Surprisingly, we find GPT-4o's compositionality inferior to the best open-source model, and we analyze the underlying reasons. Our experimental analysis reveals the limitations of VLMs in fine-grained compositional perception and reasoning, and points to areas for improvement in VLM design and training. Resources available at: https://hanghuacs.github.io/MMComposition/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。