arXiv:2605.11367cs.CV2026-05

让机器人在部分可见环境下,实时推断并更新对3D世界的信念。

3D-Belief: Embodied Belief Inference via Generative 3D World Modeling

论文配图:3D-Belief: Embodied Belief Inference via Generative 3D World Modeling
图 1 · 摘自论文原文
  • 基于3D世界建模,将不确定性显式表示在空间中
  • 支持多假设推理和动态信念更新,提升场景补全能力
  • 适用于真实世界导航,比现有方法更适应局部观测

视觉生成模型的进展展示了构建生成式世界模型的潜力。然而,现有方法大多将世界建模视为新视角合成或未来帧预测,强调视觉真实性而非具身智能体在部分可观测下的结构化不确定性。本文提出将世界建模视为3D空间中的具身信念推断。在此视角下,世界模型不仅需渲染可能看到的内容,还需随新观测持续维护和更新对未观测3D世界的信念。我们识别出关键能力:空间一致的场景记忆、多假设信念采样、序列信念更新、以及对未见区域的语义引导预测。我们构建了3D-Belief模型,从部分观测中推断显式的、可操作的3D信念,并在线更新。与以往视觉预测模型不同,3D-Belief直接在3D中表示不确定性,使具身智能体能想象合理场景补全并推理部分观测环境。我们在2D视觉质量、场景记忆与未见场景想象、对象与场景级3D想象(使用自建3D-CORE基准)以及模拟与真实世界中的物体导航任务上评估该模型。实验表明,3D-Belief在2D与3D想象质量及下游具身任务表现上均优于当前最优方法。

原文摘要 · Abstract (English)

Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame world modeling as novel-view synthesis or future-frame prediction, emphasizing visual realism rather than the structured uncertainty required by embodied agents acting under partial observability. In this work, we propose a different perspective: world modeling as embodied belief inference in 3D space. From this view, a world model should not merely render what may be seen, but maintain and update an agent's belief about the unobserved 3D world as new observations are acquired. We identify several key capabilities for such models, including spatially consistent scene memory, multi-hypothesis belief sampling, sequential belief updating, and semantically informed prediction of unseen regions. We instantiate these ideas in 3D-Belief, a generative 3D world model that infers explicit, actionable 3D beliefs from partial observations and updates them online over time. Unlike prior visual prediction models, 3D-Belief represents uncertainty directly in 3D, enabling embodied agents to imagine plausible scene completions and reason over partially observed environments. We evaluate 3D-Belief on 2D visual quality for scene memory and unobserved-scene imagination, object- and scene-level 3D imagination using our proposed 3D-CORE benchmark, and challenging object navigation tasks in both simulation and the real world. Experiments show that 3D-Belief improves 2D and 3D imagination quality and downstream embodied task performance compared to state-of-the-art methods.

3D生成具身智能信念推断世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。