arXiv:2509.25142cs.AI2025-09被引 5

视觉语言模型在复杂推理任务中表现不佳,因缺乏逐步处理视觉信息的能力。

Visual serial processing deficits explain divergences in human and VLM reasoning

  • 通过对比人类与模型在三类任务中的表现,发现序列处理需求越强,模型差距越大。
  • 人类反应时间越长(代表需更多序列处理),模型准确率下降越明显。
  • 适合关注模型认知局限、视觉推理机制的研究者阅读。

为什么视觉语言模型(VLMs)在标准基准上表现良好,却在看似简单的视觉推理任务中难以媲美人类?我们假设,关键原因在于其在视觉引导的序列处理能力上的缺陷。为验证该假设,我们在几何推理、感知计数和心理旋转三个领域设计了任务,通过调整几何概念复杂度、感知区分负荷和变换难度来控制序列处理需求。结果显示:随着任务对序列处理的要求上升,模型准确率显著下降,且与人类反应时间(作为序列处理负荷的代理指标)呈强相关。当任务需要组合概念、计数物体或执行心理变换时,模型与人类的性能差距系统性扩大。这支持了我们的假设:当前VLMs在视觉引导的序列推理能力上的局限,是其与人类表现差异的根本瓶颈。

原文摘要 · Abstract (English)

Why do Vision Language Models (VLMs), despite success on standard benchmarks, often fail to match human performance on surprisingly simple visual reasoning tasks? While the underlying computational principles are still debated, we hypothesize that a crucial factor is a deficit in visually-grounded serial processing. To test this hypothesis, we compared human and VLM performance across tasks designed to vary serial processing demands in three distinct domains: geometric reasoning, perceptual enumeration, and mental rotation. Tasks within each domain varied serial processing load by manipulating factors such as geometric concept complexity, perceptual individuation load, and transformation difficulty. Across all domains, our results revealed a consistent pattern: decreased VLM accuracy was strongly correlated with increased human reaction time (used as a proxy for serial processing load). As tasks require more demanding serial processing -- whether composing concepts, enumerating items, or performing mental transformations -- the VLM-human performance gap widens reliably. These findings support our hypothesis, indicating that limitations in serial, visually grounded reasoning represent a fundamental bottleneck that distinguishes current VLMs from humans.

视觉推理模型局限认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。