让视觉语言模型通过世界模型实现3D空间推理,无需微调即可提升表现。
MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
- 用视频扩散模型构建可控世界模型,辅助视觉语言模型生成多视角视图。
- 在SAT基准上平均提升7.7%性能,且优于强化学习训练的模型。
- 适用于需要3D空间理解的导航与操作任务,适合研究具身智能的学者。
三维空间中的空间推理是人类认知的核心,对具身任务如导航和操作至关重要。然而,当前最先进的视觉语言模型(VLM)常难以预测自我中心运动后的场景变化:它们仅处理二维图像,缺乏对三维动态的内部建模能力。为此,我们提出MindJourney,一种测试时扩展框架,通过将VLM与基于视频扩散的可控世界模型结合,赋予其缺失的三维动态建模能力。VLM迭代绘制简明相机轨迹,世界模型在每一步合成对应视角;VLM随后基于交互式探索获得的多视角证据进行推理。无需任何微调,MindJourney在代表性空间推理基准SAT上平均提升7.7%性能,表明将VLM与世界模型结合进行测试时扩展,是一种简单、可即插即用的鲁棒3D推理路径。同时,该方法也优于通过强化学习训练的测试时推理VLM,展示了利用世界模型实现测试时扩展的巨大潜力。
原文摘要 · Abstract (English)
Spatial reasoning in 3D space is central to human cognition and indispensable for embodied tasks such as navigation and manipulation. However, state-of-the-art vision-language models (VLMs) struggle frequently with tasks as simple as anticipating how a scene will look after an egocentric motion: they perceive 2D images but lack an internal model of 3D dynamics. We therefore propose MindJourney, a test-time scaling framework that grants a VLM with this missing capability by coupling it to a controllable world model based on video diffusion. The VLM iteratively sketches a concise camera trajectory, while the world model synthesizes the corresponding view at each step. The VLM then reasons over this multi-view evidence gathered during the interactive exploration. Without any fine-tuning, our MindJourney achieves over an average 7.7% performance boost on the representative spatial reasoning benchmark SAT, showing that pairing VLMs with world models for test-time scaling offers a simple, plug-and-play route to robust 3D reasoning. Meanwhile, our method also improves upon the test-time inference VLMs trained through reinforcement learning, which demonstrates the potential of our method that utilizes world models for test-time scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。