大模型解空间谜题时会形成心理图像,甚至'想象'羊。
Do multimodal models imagine electric sheep?

- 通过预测动作序列训练模型,使其激活值编码中间状态视觉信息。
- 引入16个视觉标记后,解题成功率从83%提升至89%,尤其在复杂任务上。
- 揭示无需显式视觉监督也能生成心理图像,适合研究具身智能者。
我们发现,大型多模态模型在解决空间谜题时会产生心理图像,且在解羊形谜题时会‘想象’羊。通过微调Qwen3.5 VLM以解决十二类多样化的视觉推理任务——包括七巧板、拼图、推箱子、三维心理旋转和交通堵塞等——这些任务要求理解几何、空间关系及动作后果。通过监督模型从初始状态预测开环动作序列,我们发现模型每一步后的激活值均编码了关于中间状态的有意义视觉信息。这表明,在缺乏任何显式视觉监督的情况下,学习正确动作选择的过程中,会自然产生一个不完善的视觉世界模型。基于此观察,我们提出了两种方法来增强并利用模型形成的内在心理图像。结果显示,每步仅引入16个视觉标记,平均解题率便从83%提升至89%,尤其在拼图和三维心理旋转等推理密集型任务中提升显著。
原文摘要 · Abstract (English)
Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. We fine-tune a Qwen3.5 VLM to solve twelve diverse visual reasoning tasks -- including tangram, jigsaw, sokoban, 3D mental rotation, and rush hour -- that require understanding geometry, spatial relationships, and the consequences of actions. By supervising the model to predict the open-loop sequence of actions to solve a puzzle from an initial state, we show that the model's activations after each action encode meaningful visual information about the intermediate state. This finding suggests that an imperfect visual world model begins to form as a byproduct of learning to select correct actions, in the absence of any explicit visual supervision. Building on this observation, we propose two ways to sharpen and use the mental images formed by the model. We find that integrating as few as sixteen visual tokens per step into the chain of thought improves the average solve rate from 83% to 89%, with particularly strong gains on reasoning-heavy tasks such as jigsaw and 3D mental rotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。