arXiv:2603.06445cs.CV2026-03中稿 · ECCV被引 1

让模型通过想象模拟探索,解答空间‘如果…会怎样’问题。

What if? Emulative Simulation with World Models for Situated Reasoning

  • 用全景视频构建心理探索轨迹数据集,支持无主动探索的推理。
  • 世界模型在数据集上表现优异,想象力显著提升问答准确率。
  • 适合研究具身智能、视觉推理与认知模拟的研究者使用。

具身推理常依赖主动探索,但在机器人物理限制或视障用户安全考虑下,主动探索难以实现。仅基于有限观察,智能体能否在脑海中模拟通向目标状态的未来轨迹,并回答空间类‘如果…会怎样’问题?我们提出WanderDream,首个面向心理模拟探索的大规模数据集,使模型无需实际探索即可进行推理。WanderDream-Gen包含15.8万段全景视频,覆盖HM3D、ScanNet++及真实场景中的1088个真实环境,描绘从当前视角到目标情境的想象路径。WanderDream-QA包含15.8万个问答对,涵盖每条轨迹的起始状态、路径与终点状态,全面评估基于探索的推理能力。大量实验表明:(1) 心理探索对具身推理至关重要;(2) 世界模型在WanderDream-Gen上表现出色;(3) 想象力显著提升在WanderDream-QA上的推理性能;(4) WanderDream数据在真实场景中具有显著迁移能力。

原文摘要 · Abstract (English)

Situated reasoning often relies on active exploration, yet in many real-world scenarios such exploration is infeasible due to physical constraints of robots or safety concerns of visually impaired users. Given only a limited observation, can an agent mentally simulate a future trajectory toward a target situation and answer spatial what-if questions? We introduce WanderDream, the first large-scale dataset designed for the emulative simulation of mental exploration, enabling models to reason without active exploration. WanderDream-Gen comprises 15.8K panoramic videos across 1,088 real scenes from HM3D, ScanNet++, and real-world captures, depicting imagined trajectories from current viewpoints to target situations. WanderDream-QA contains 158K question-answer pairs, covering starting states, paths, and end states along each trajectory to comprehensively evaluate exploration-based reasoning. Extensive experiments with world models and MLLMs demonstrate (1) that mental exploration is essential for situated reasoning, (2) that world models achieve compelling performance on WanderDream-Gen, (3) that imagination substantially facilitates reasoning on WanderDream-QA, and (4) that WanderDream data exhibit remarkable transferability to real-world scenarios.

具身推理心理模拟世界模型空间问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。