用物理力控生成视频,无需3D模拟器也能真实响应推拉风等力。
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
- 通过力提示让视频模型响应点力与风场等物理控制信号。
- 仅用1.5万样本训练,对多种物体材质和场景都能真实模拟受力效果。
- 适合做物理交互式视频生成的研究者与开发者使用。
近期视频生成模型的发展激发了对能够模拟真实环境的世界模型的兴趣。尽管导航已广泛研究,但能模仿真实世界力的物理交互仍鲜有探索。本文提出力提示机制,使用户可通过局部点力(如轻戳植物)或全局风场(如风吹布料)与图像互动。我们证明,仅利用预训练模型中的视觉与运动先验,无需3D资产或物理模拟器,视频即可真实响应物理控制信号。主要挑战在于高质量配对力-视频数据的获取:真实数据难采集,合成数据受限于视觉质量与领域多样性。关键发现是,模型在仅少量物体示范下,经由Blender合成视频微调后,仍能显著泛化至不同几何、场景与材料。我们进一步分析其泛化来源,发现视觉多样性与特定文本关键词是两个核心因素。方法仅需约1.5万训练样本,在四块A100 GPU上训练一天,便在力遵循度与物理真实性上优于现有方法,推动世界模型向真实物理交互迈进。所有数据集、代码、权重及交互演示均已开源。
原文摘要 · Abstract (English)
Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has been well-explored, physically meaningful interactions that mimic real-world forces remain largely understudied. In this work, we investigate using physical forces as a control signal for video generation and propose force prompts which enable users to interact with images through both localized point forces, such as poking a plant, and global wind force fields, such as wind blowing on fabric. We demonstrate that these force prompts can enable videos to respond realistically to physical control signals by leveraging the visual and motion prior in the original pretrained model, without using any 3D asset or physics simulator at inference. The primary challenge of force prompting is the difficulty in obtaining high quality paired force-video training data, both in the real world due to the difficulty of obtaining force signals, and in synthetic data due to limitations in the visual quality and domain diversity of physics simulators. Our key finding is that video generation models can generalize remarkably well when adapted to follow physical force conditioning from videos synthesized by Blender, even with limited demonstrations of few objects. Our method can generate videos which simulate forces across diverse geometries, settings, and materials. We also try to understand the source of this generalization and perform ablations that reveal two key elements: visual diversity and the use of specific text keywords during training. Our approach is trained on only around 15k training examples for a single day on four A100 GPUs, and outperforms existing methods on force adherence and physics realism, bringing world models closer to real-world physics interactions. We release all datasets, code, weights, and interactive video demos at our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。