arXiv:2601.05848cs.CVcs.AI2026-01被引 10

用物理力向量指导视频模型完成复杂任务,无需外部物理引擎。

Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals

  • 通过力向量和动态过程定义目标,模拟人类对物理任务的理解。
  • 在真实场景中零样本泛化,实现工具操作与多物体因果链任务。
  • 适合机器人规划、物理仿真研究者,可直接用于交互式视频生成。

近期视频生成进展推动了“世界模型”的发展,使其能为机器人和规划任务模拟未来可能。然而,精确设定目标仍具挑战:文本指令过于抽象,难以捕捉物理细节;目标图像又常因动态任务而不可行。为此,我们提出 Goal Force 框架,允许用户通过显式的力向量和中间动态过程定义目标,模仿人类对物理任务的思维模式。我们在合成因果基元数据集(如弹性碰撞、倒多米诺骨牌)上训练视频生成模型,使其能够随时间空间传播力。尽管仅在简单物理数据上训练,该模型在复杂真实场景中展现出显著零样本泛化能力,涵盖工具操控和多物体因果链。结果表明,通过将视频生成锚定于基本物理相互作用,模型可自发成为隐式神经物理模拟器,实现无需依赖外部引擎的精准、物理感知规划。项目已开源数据集、代码、模型权重及交互式视频演示。

原文摘要 · Abstract (English)

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text instructions are often too abstract to capture physical nuances, while target images are frequently infeasible to specify for dynamic tasks. To address this, we introduce Goal Force, a novel framework that allows users to define goals via explicit force vectors and intermediate dynamics, mirroring how humans conceptualize physical tasks. We train a video generation model on a curated dataset of synthetic causal primitives-such as elastic collisions and falling dominos-teaching it to propagate forces through time and space. Despite being trained on simple physics data, our model exhibits remarkable zero-shot generalization to complex, real-world scenarios, including tool manipulation and multi-object causal chains. Our results suggest that by grounding video generation in fundamental physical interactions, models can emerge as implicit neural physics simulators, enabling precise, physics-aware planning without reliance on external engines. We release all datasets, code, model weights, and interactive video demos at our project page.

视频生成物理模拟机器人规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。