25个视觉语言模型中24个无法正确画出视错觉中的水平线,暴露了空间推理短板。
Constructive Apraxia: An Unexpected Limit of Instructible Vision-Language Models and Analog for Human Cognitive Disorders
- 用视错觉测试模型空间理解能力,发现多数模型误将直线按透视倾斜
- 24/25模型失败,与大脑顶叶损伤患者的构图障碍高度相似
- 为研究人类认知缺陷提供新计算模型,适合关注AI认知局限的研究者
本研究揭示了可指令化视觉语言模型(VLMs)与人类认知障碍——构图失用症之间的意外关联。我们测试了25个前沿VLMs,包括GPT-4 Vision、DALL-E 3和Midjourney v5,要求其生成庞佐错觉图像,该任务需基础空间推理,常用于临床评估构图失用症。惊人的是,24/25个模型未能正确绘制两条水平线,反而沿背景透视方向倾斜,表现出与顶叶损伤患者相同的构图缺陷:虽视觉感知和运动能力完好,却无法准确复制或构建简单图形。这表明当前VLM尽管在其他领域表现优异,仍缺乏类似构图失用症患者的空间推理能力。该缺陷为研究空间认知障碍提供了新的计算模型,并凸显了改进VLM架构与训练方法的紧迫性。
原文摘要 · Abstract (English)
This study reveals an unexpected parallel between instructible vision-language models (VLMs) and human cognitive disorders, specifically constructive apraxia. We tested 25 state-of-the-art VLMs, including GPT-4 Vision, DALL-E 3, and Midjourney v5, on their ability to generate images of the Ponzo illusion, a task that requires basic spatial reasoning and is often used in clinical assessments of constructive apraxia. Remarkably, 24 out of 25 models failed to correctly render two horizontal lines against a perspective background, mirroring the deficits seen in patients with parietal lobe damage. The models consistently misinterpreted spatial instructions, producing tilted or misaligned lines that followed the perspective of the background rather than remaining horizontal. This behavior is strikingly similar to how apraxia patients struggle to copy or construct simple figures despite intact visual perception and motor skills. Our findings suggest that current VLMs, despite their advanced capabilities in other domains, lack fundamental spatial reasoning abilities akin to those impaired in constructive apraxia. This limitation in AI systems provides a novel computational model for studying spatial cognition deficits and highlights a critical area for improvement in VLM architecture and training methodologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。