用认知科学的绑定问题解释视觉语言模型为何会失败
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem
- 从绑定问题出发,分析模型在多物体任务中的表现
- 模型在计数、定位等任务上表现差,与人类快速处理机制相似
- 揭示了先进模型在复杂推理中的本质局限,适合研究者参考
近期研究表明,最先进的视觉语言模型(VLMs)在描述和生成复杂自然图像方面表现优异,但在基本的多物体推理任务(如计数、定位和简单视觉类比)上却出现令人惊讶的失败,而人类在这些任务上几乎完美。为理解这一看似矛盾的表现模式,本文借鉴认知科学与神经科学中的绑定问题理论——即当共享表征资源需表示多个独立实体时,必须通过串行处理避免干扰。研究发现,许多先进VLM的失败可归因于绑定问题,且其失败模式与人脑快速前馈处理的局限性高度相似。
原文摘要 · Abstract (English)
Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language models and text-to-image models. These models are able to describe and generate a diverse array of complex, naturalistic images, yet they exhibit surprising failures on basic multi-object reasoning tasks -- such as counting, localization, and simple forms of visual analogy -- that humans perform with near perfect accuracy. To better understand this puzzling pattern of successes and failures, we turn to theoretical accounts of the binding problem in cognitive science and neuroscience, a fundamental problem that arises when a shared set of representational resources must be used to represent distinct entities (e.g., to represent multiple objects in an image), necessitating the use of serial processing to avoid interference. We find that many of the puzzling failures of state-of-the-art VLMs can be explained as arising due to the binding problem, and that these failure modes are strikingly similar to the limitations exhibited by rapid, feedforward processing in the human brain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。