测试视觉语言模型能否自动关联相似视觉线索,发现能力不足且存在明显差距。
VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues
- 构建9个子任务、3000+案例的基准测试,评估模型视觉线索关联能力
- 12个模型表现参差,普遍在跨图匹配中出现显著性能下降
- 适合研究多模态理解、视觉推理与模型可解释性的学者参考
在日常生活中,根据视觉线索(如外貌特征)识别同一人出现在多张照片中是一项关键能力,即使不知道其身份。尽管视觉语言模型(VLMs)具备丰富知识,但它们是否能完成这一基础任务仍不清楚。为此,我们提出 extbf{VLM2-Bench},一个用于评估 VLMs 是否能视觉链接匹配线索的基准,包含9个子任务和超过3,000个测试用例。对12个主流VLMs进行全面评估,并进一步分析不同语言侧与视觉侧提示方法的影响,得出8项关键发现。研究揭示了模型在视觉线索关联上的核心挑战,凸显显著性能差距。基于此,我们建议:(i) 提升底层视觉能力以增强适应性并减少对先验知识依赖;(ii) 明确语言推理在视觉主导任务中的整合原则,避免引入不必要的偏差;(iii) 推动视觉-文本训练范式向促进模型独立构建与推断视觉线索关系的能力转变。
原文摘要 · Abstract (English)
Visually linking matching cues is a crucial ability in daily life, such as identifying the same person in multiple photos based on their cues, even without knowing who they are. Despite the extensive knowledge that vision-language models (VLMs) possess, it remains largely unexplored whether they are capable of performing this fundamental task. To address this, we introduce \textbf{VLM2-Bench}, a benchmark designed to assess whether VLMs can Visually Link Matching cues, with 9 subtasks and over 3,000 test cases. Comprehensive evaluation across twelve VLMs, along with further analysis of various language-side and vision-side prompting methods, leads to a total of eight key findings. We identify critical challenges in models' ability to link visual cues, highlighting a significant performance gap. Based on these insights, we advocate for (i) enhancing core visual capabilities to improve adaptability and reduce reliance on prior knowledge, (ii) establishing clearer principles for integrating language-based reasoning in vision-centric tasks to prevent unnecessary biases, and (iii) shifting vision-text training paradigms toward fostering models' ability to independently structure and infer relationships among visual cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。