arXiv:2606.00439cs.CV2026-06

用可控制的物理世界模型,从视频中自动发现物体及其运动规律。

Physical Object Understanding with a Physically Controllable World Model

论文配图:Physical Object Understanding with a Physically Controllable World Model
图 1 · 摘自论文原文
  • 基于自回归序列建模,学习视觉变量的条件概率分布。
  • 生成多组未来场景,通过运动关联性识别物体与部件。
  • 能3D操控物体,计算物理关系,支持如视觉积木等应用。

视觉智能的核心挑战之一是从原始视频中学习场景的物理结构:区域如何形成物体,以及它们相互作用的规律。解决这些问题需要能够从部分观测中推断世界分布状态的世界模型,而现有架构尚不具备此类能力。我们提出一类新的概率世界模型,可对任意视觉变量(如外观和动态)进行概率估计,条件依赖于其他变量。通过自回归序列建模,该模型可高效训练,并从中涌现出丰富的物体理解。首先,模型通过序列推断生成多个可能的未来世界状态,捕捉物体运动的物理规律;其次,分析这些未来状态间的运动相关性,实现物体及可动部件的自动识别;在发现物体后,模型可对物体进行3D操控;最后,利用模型推导物体间的物理关系,支持如视觉积木(Visual Jenga)等应用。

原文摘要 · Abstract (English)

A central challenge in visual intelligence is learning the physical structure of scenes from raw videos: how regions form objects and the laws that govern their interactions. Solving these tasks requires world models capable of inferring distributional states of the world from partial observations - capabilities that current architectures do not provide. We introduce a new class of probabilistic world models that support estimation of the probability of any visual variable, such as appearance and dynamics, conditioned on any other variables. Here, we identify that these models can be trained efficiently with autoregressive sequence modeling, yielding world models from which rich object understanding emerges. First, we demonstrate that our model captures the physical laws governing how objects move by generating multiple plausible future states of the world through sequential inference. Then, by analyzing motion correlations across these futures, we extract objects and articulated object subparts. Having discovered these objects, we show that our world model can manipulate them in 3D. Finally, we demonstrate how physical relationships between objects can be computed from the world model, enabling applications such as Visual Jenga.

世界模型物理理解物体识别3D操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。