用强化学习让视频生成模型懂物理,提升真实感。
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
- 通过编码输入图像中的物理信息作为条件引导生成
- 利用人类反馈优化物理表征,提升生成视频的物理合理性
- 可通用适配多种物理场景,适合需要真实物理行为的应用
当前视频生成模型虽能生成视觉逼真的视频,但常违背物理规律,难以作为可信的‘世界模型’。为此,我们提出PhysMaster,将物理知识建模为表征,用于引导视频生成以增强物理感知能力。针对图像到视频任务,模型需从单帧图像预测符合物理规律的动态过程。由于输入图像包含物体相对位置和潜在相互作用等物理先验,我们设计了PhysEncoder来提取这些物理信息,并作为额外条件注入生成流程。由于缺乏对物理表现的显式监督,我们采用基于人类反馈的强化学习方法,结合直接偏好优化(DPO)端到端优化物理表征。实验表明,PhysMaster在简单代理任务上有效提升物理一致性,并具备在多样物理场景中的泛化能力,证明其可通过表示学习统一处理各类物理过程,是一种通用且可插拔的物理感知视频生成方案。
原文摘要 · Abstract (English)
Video generation models nowadays are capable of generating visually realistic videos, but often fail to adhere to physical laws, limiting their ability to generate physically plausible videos and serve as ''world models''. To address this issue, we propose PhysMaster, which captures physical knowledge as a representation for guiding video generation models to enhance their physics-awareness. Specifically, PhysMaster is based on the image-to-video task where the model is expected to predict physically plausible dynamics from the input image. Since the input image provides physical priors like relative positions and potential interactions of objects in the scenario, we devise PhysEncoder to encode physical information from it as an extra condition to inject physical knowledge into the video generation process. The lack of proper supervision on the model's physical performance beyond mere appearance motivates PhysEncoder to apply reinforcement learning with human feedback to physical representation learning, which leverages feedback from generation models to optimize physical representations with Direct Preference Optimization (DPO) in an end-to-end manner. PhysMaster provides a feasible solution for improving physics-awareness of PhysEncoder and thus of video generation, proving its ability on a simple proxy task and generalizability to wide-ranging physical scenarios. This implies that our PhysMaster, which unifies solutions for various physical processes via representation learning in the reinforcement learning paradigm, can act as a generic and plug-in solution for physics-aware video generation and broader applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。