arXiv:2603.00110cs.RO2026-03被引 3

用视频模型学物理,让机器人学会抓取透明物体。

Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation

  • 把预训练视频模型当物理模拟器,统一视觉与动作的连续表示。
  • 在真实场景中表现媲美大规模动作预训练模型,提升13.8%成功率。
  • 适合想低成本构建通用机械臂控制系统的研究人员。

由于大规模机器人数据稀缺,现有方法尝试复用其他模态的基础模型进行策略学习。本文提出PhysGen(从预训练视频生成模型中学习物理),一种可扩展的连续与序列世界交互框架,利用自回归视频生成模型解决机器人操作任务。通过将预训练视频模型视为物理模拟器,PhysGen建模外部环境与机器人动作之间的动态交互。引入多模态连续表征,将视频与动作统一为共享的物理标记,弥合离散视频生成与连续机器人控制之间的鸿沟。该方法实现隐式物理知识(如物体恒常性与动力学)从视频预训练到下游操作任务的无缝迁移。为确保高效收敛,引入因果掩码、逆运动学、前瞻多标记预测(L-MTP)及键值缓存机制。在Libero和ManiSkill基准上的实验表明,PhysGen持续优于强基线,分别超越OpenVLA和WorldVLA 13.8%和8.8%。值得注意的是,在真实场景中,PhysGen在无需特定动作预训练的情况下,性能达到大型动作预训练模型π₀的水平,尤其在抓取透明物体等复杂物理任务中表现优异。结果验证了从预训练视频生成器中提取物理直觉以促进可泛化机器人操作的潜力。

原文摘要 · Abstract (English)

The scarcity of large-scale robotic data has motivated the repurposing of foundation models from other modalities for policy learning. In this work, we introduce PhysGen (Learning Physics from Pretrained Video Generation Models), a scalable continuous and sequential world interaction framework that leverages autoregressive video generation to solve robotic manipulation tasks. By treating the pretrained video model as a proxy for a physics simulator, PhysGen models the dynamic interplay between the external environment and robot actions. We introduce a multimodal continuous representation that unifies video and action into shared physical tokens, bridging the gap between discrete video generation and continuous robotic control. This approach enables the seamless transfer of implicit physical knowledge-such as object permanence and dynamics-from video pretraining to downstream manipulation.To ensure efficient convergence, we incorporate causal masking, inverse kinematics, Lookahead Multi-Token Prediction (L-MTP), and key-value (KV) caching. Experimental results on the Libero and ManiSkill benchmarks demonstrate that PhysGen consistently outperforms robust baselines, surpassing OpenVLA and WorldVLA by margins of 13.8% and 8.8%, respectively. Notably, in real-world scenarios, PhysGen matches the performance of large-scale action-pretrained models like $π_0$ without requiring prior action-specific pretraining, demonstrating superior capability in physically complex tasks such as grasping transparent objects. These findings validate the potential of extracting physical intuition from pretrained video generators to facilitate generalizable robotic manipulation.

机器人视频生成物理建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。