arXiv:2509.15271cs.CVcs.AI2025-09中稿 · ICASSP 2026被引 5

大视觉模型能像人一样进行空间旋转推理,揭示其认知机制。

Large Vision Models Can Solve Mental Rotation Problems

  • 用分层探测法评估多种视觉模型的空间推理能力
  • 自监督模型在几何结构捕捉上优于监督模型,中间层表现更佳
  • 结果与人类反应时间模式相似,适合研究视觉认知机制

空间旋转是人类空间推理的关键测试,对理解感知如何支持认知具有重要意义。尽管现代视觉变换器取得显著进展,但它们是否具备类似能力仍不明确。本文系统评估了ViT、CLIP、DINOv2和DINOv3在一系列心理旋转任务中的表现,涵盖从谢帕德与梅茨勒使用的简单积木结构,到复杂积木图形、三类文本以及逼真物体。通过逐层探测模型表示,我们分析其成功机制。发现:i)自监督视觉变换器比监督模型更好地捕捉几何结构;ii)中间层的表现优于最终层;iii)任务难度随旋转复杂度和遮挡程度增加而上升,与人类反应时间模式一致,表明嵌入空间表示存在类似约束。

原文摘要 · Abstract (English)

Mental rotation is a key test of spatial reasoning in humans and has been central to understanding how perception supports cognition. Despite the success of modern vision transformers, it is still unclear how well these models develop similar abilities. In this work, we present a systematic evaluation of ViT, CLIP, DINOv2, and DINOv3 across a range of mental-rotation tasks, from simple block structures similar to those used by Shepard and Metzler to study human cognition, to more complex block figures, three types of text, and photo-realistic objects. By probing model representations layer by layer, we examine where and how these networks succeed. We find that i) self-supervised ViTs capture geometric structure better than supervised ViTs; ii) intermediate layers perform better than final layers; iii) task difficulty increases with rotation complexity and occlusion, mirroring human reaction times and suggesting similar constraints in embedding space representations.

空间推理视觉模型认知机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。