用自然语言描述物理世界变化,让模型先推理再生成视频。
PhiZero: A World Model Built Around Physical Language

- 用自监督学习从真实视频中提取物理语言表示
- 在多个任务上实现物理一致的未来预测与交互建模
- 适合需要精细动作控制和零样本运动迁移的研究者
我们提出PhiZero,一种基于物理语言的物理世界模型。物理语言是一种紧凑的离散世界状态转移表示。现有物理世界模型通常直接在像素空间预测未来视频,使底层世界动态隐含于高维视觉预测器中。受人类从视觉经验中抽象预测结构并以自然语言组织进行显式推理的启发,我们通过自监督学习从真实视频中习得物理语言,并利用它显式推理物理世界的演化过程。因此,PhiZero采用先推理后渲染的范式:先推断未来的物理语言序列,再将其转化为视频。在生成与理解基准上的大量实验验证了其建模物理一致性世界演化的潜力。我们还展示了其在真实感、交互式世界建模、细粒度动作条件模拟及零样本运动迁移方面的应用前景。
原文摘要 · Abstract (English)
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Motivated by humans' ability to abstract predictive structure from visual experience and organize it in natural language for explicit reasoning, we learn physical language from in-the-wild videos through self-supervision and use it to explicitly reason about how the physical world evolves. Accordingly, PhiZero adopts a reason-then-render paradigm: it first infers future world evolution as a physical-language sequence and then renders the inferred transitions into videos. Extensive experiments across generation and understanding benchmarks validate the ability of PhiZero to model physically coherent world evolution. We further show its potential for realistic and interactive world modeling, fine-grained action-conditioned simulation, and zero-shot motion transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。