让机器人操作视频既真实又符合物理规律,还能精准控制动作。
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
- 用300万条带物理标注的操作视频训练140亿参数扩散模型
- 在真实与合成场景中均超越Sora和Veo的物理合理性与轨迹一致性
- 提出新评估基准EZSbench,专测未见任务的零样本泛化能力
基于视频的世界模型为具身模拟与规划提供了强大范式,但当前先进模型常生成违反物理规律的操作(如物体穿透、反重力运动),因其训练数据为通用视觉数据,且目标函数基于似然性而忽略物理定律。本文提出ABot-PhysWorld,一个140亿参数的扩散变压器模型,可生成视觉真实、物理合理且动作可控的视频。该模型基于包含三百万条操作片段的精选数据集,采用新型基于直接偏好优化(DPO)的后训练框架,通过解耦判别器抑制非物理行为,同时保持视觉质量。并行上下文模块实现精确空间动作注入,支持跨具身控制。为更好评估泛化能力,提出首个独立于训练的具身零样本评估基准EZSbench,结合真实与合成的未见机器人-任务-场景组合,并采用解耦协议分别评估物理真实性与动作对齐度。ABot-PhysWorld在PBench和EZSbench上达到新最优表现,优于Veo 3.1与Sora v2 Pro。我们将公开发布EZSbench以推动具身视频生成领域的标准化评估。
原文摘要 · Abstract (English)
Video-based world models offer a powerful paradigm for embodied simulation and planning, yet state-of-the-art models often generate physically implausible manipulations - such as object penetration and anti-gravity motion - due to training on generic visual data and likelihood-based objectives that ignore physical laws. We present ABot-PhysWorld, a 14B Diffusion Transformer model that generates visually realistic, physically plausible, and action-controllable videos. Built on a curated dataset of three million manipulation clips with physics-aware annotation, it uses a novel DPO-based post-training framework with decoupled discriminators to suppress unphysical behaviors while preserving visual quality. A parallel context block enables precise spatial action injection for cross-embodiment control. To better evaluate generalization, we introduce EZSbench, the first training-independent embodied zero-shot benchmark combining real and synthetic unseen robot-task-scene combinations. It employs a decoupled protocol to separately assess physical realism and action alignment. ABot-PhysWorld achieves new state-of-the-art performance on PBench and EZSbench, surpassing Veo 3.1 and Sora v2 Pro in physical plausibility and trajectory consistency. We will release EZSbench to promote standardized evaluation in embodied video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。