arXiv:2606.28128cs.CVcs.AI2026-06被引 3

让机器人操作视频更符合物理规律,提升仿真可靠性。

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

论文配图:PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
图 1 · 摘自论文原文
  • 通过像素与语义双重监督,强化运动轨迹和物体交互的物理一致性。
  • 在R-Bench上使基线模型性能提升最高达22.3%,闭环成功率从16%升至24%。
  • 适合需要高可信度世界模拟的机器人任务,如动作规划与策略训练。

视频生成模型已成为具身世界模拟的有前途范式。然而,通用视频生成器及针对机器人数据微调的模型仍会产生不符合物理规律的操作,如运动轨迹不连续、机器人-物体交互不一致,限制其作为世界模拟器的可靠性。通过大量实验发现,这类物理不稳定性主要源于移动物体形变及交互实体间时空关联不合理,尤其在接触阶段。基于此,我们提出PhysisForcing,一种可扩展的训练框架,通过联合优化像素级与语义级特征,在物理信息区域施加聚焦监督。该框架包含像素级轨迹对齐损失(使用参考点轨迹监督DiT特征)和语义级关系对齐损失(将DiT特征与冻结视频理解编码器提取的区域间关系对齐)。在R-Bench、PAI-Bench和EZS-Bench上的实验表明,PhysisForcing持续优于强基线模型,在R-Bench上使Wan2.2-I2V-A14B和Cosmos3-Nano基线分别提升22.3%和9.2%(较纯微调提升7.1%和3.7%),其中Cosmos3-Nano版本取得最优综合得分。除生成外,作为WorldArena行动规划协议下的世界模型,其将闭环成功率从16.0%提升至24.0%,并进一步提高下游策略成功率,表明物理对齐的视频模型能生成更强的机器人操作表征。

原文摘要 · Abstract (English)

Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models can still produce physically implausible manipulations, including discontinuous motion trajectories and inconsistent robot-object interactions, which limits their reliability as world simulators. Through extensive experiments, we find that such physical instability mainly arises from two factors: deformation of moving objects and implausible spatio-temporal correlations among interacting entities, particularly during contact. Building on this observation, we propose PhysisForcing, a scalable training framework that strengthens physical consistency by focusing supervision on physics-informative regions through joint optimization of pixel-level and semantic-level features. The framework consists of a pixel-level trajectory alignment loss, which supervises DiT features using reference point trajectories, and a semantic-level relational alignment loss, which aligns DiT features with inter-region relations extracted from a frozen video understanding encoder. Extensive experiments on R-Bench, PAI-Bench, and EZS-Bench show that PhysisForcing consistently improves embodied video generation over strong baselines, improving the Wan2.2-I2V-A14B and Cosmos3-Nano base models on R-Bench by 22.3\% and 9.2\% (7.1\% and 3.7\% over vanilla finetuning), with the Cosmos3-Nano variant attaining the best overall score. Beyond generation, as a world model under the WorldArena action-planner protocol it raises the closed-loop success rate from 16.0\% to 24.0\% and further improves downstream policy success, indicating that physically aligned video models yield stronger representations for robotic manipulation.

世界模型机器人操作物理约束视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。