用视觉输入模拟物体在真实世界中的动态变化,支持部分观测与不完整操作信号。
Perceptual 3D Simulation With Physical World Modeling

- 基于感知推断多模态场景变量分布,实现不确定性下的预测
- 融合几何约束与持续记忆,支持在线更新和长期一致性
- 适用于新视角生成、物体操作等多样3D任务,适合机器人与虚拟场景应用
从图像中预测经过特定3D变换后场景的演化是视觉、图形学和机器人领域的核心目标。然而,真实系统受限于感知输入和局部动作,其信息本质上是部分且不完整的。本文提出P3Sim,一个能在部分观测和不完整3D变换信号下模拟未来场景状态的物理世界建模系统。P3Sim由三个相互作用的模块构成:可学习的物理世界模型、几何条件模块和持久化场景记忆。世界模型将感知视为对多模态场景变量的概率推理,可对任意组合变量提供条件分布预测。几何条件模块在推理时提供部分3D变换信号以指导模型。持久化场景记忆通过时间累积预测,实现在线更新与不确定性下的稳定性。该设计结合了数据驱动的灵活性与显式几何结构的归纳偏置,使P3Sim在多种3D变换任务(如新视角合成、物体操作、动态场景预测)中表现出强泛化能力,推动通用3D场景理解与变换的发展。
原文摘要 · Abstract (English)
Predicting how a scene will evolve after a desired 3D transformation from images is a central goal in vision, graphics, and robotics. Yet unlike ideal simulators with full access to 3D geometry and dynamics, real world systems must rely on perceptual inputs and local actions that are inherently partial and incomplete. In this work, we present P3Sim, a physical world modeling system that simulates future scene states under both partial observations and incomplete 3D transformation signals. P3Sim is composed of three interacting components: a learned physical world model, a geometric conditioning module, and a persistent scene memory. The world model interprets perception as probabilistic inference over multimodal scene variables, providing predictions of the distributions of any scene variable conditioned on any combination of others. The geometric conditioning module provides a partial 3D transform signal for conditioning the world model at inference time. The persistent scene memory integrates predictions over time, enabling online updates and consistency under uncertainty. By combining learned inference with explicit geometric structure, P3Sim balances data-driven flexibility with built-in inductive bias. This design yields a flexible perceptual simulator that generalizes across diverse 3D transformation tasks, such as novel view synthesis, object manipulation, and dynamic scene prediction, advancing toward general purpose 3D scene understanding and transformation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。