用符号推理+强化学习让扩散模型生成更符合物理规律的视频
Reasoning Physical Video Generation with Diffusion Timestep Tokens via Reinforcement Learning
- 通过递归视觉标记提取丢失的视觉属性,实现符号化表示
- 两阶段训练:监督微调传递知识,强化学习优化物理一致性
- 可动态调整生成视频的物理特性,适合需要真实物理行为的场景
尽管视频生成领域取得进展,但生成符合物理规律的视频仍是重大挑战。传统扩散方法因依赖数据驱动近似,难以泛化到未见物理条件(如速度)。为此,我们提出结合符号推理与强化学习,以确保视频生成的物理一致性。首先引入扩散时间步标记器(DDT),通过恢复扩散过程中丢失的视觉属性,学习离散的递归视觉标记。这些标记使大型语言模型能够进行符号推理。基于此,我们提出Phys-AR框架,包含两个阶段:第一阶段采用监督微调转移符号知识,第二阶段通过基于物理条件的奖励函数,应用强化学习优化模型的推理能力。该方法使模型能动态调整并改进生成视频的物理属性,确保其遵循物理定律。实验结果表明,PhysAR能生成物理一致的视频。
原文摘要 · Abstract (English)
Despite recent progress in video generation, producing videos that adhere to physical laws remains a significant challenge. Traditional diffusion-based methods struggle to extrapolate to unseen physical conditions (eg, velocity) due to their reliance on data-driven approximations. To address this, we propose to integrate symbolic reasoning and reinforcement learning to enforce physical consistency in video generation. We first introduce the Diffusion Timestep Tokenizer (DDT), which learns discrete, recursive visual tokens by recovering visual attributes lost during the diffusion process. The recursive visual tokens enable symbolic reasoning by a large language model. Based on it, we propose the Phys-AR framework, which consists of two stages: The first stage uses supervised fine-tuning to transfer symbolic knowledge, while the second stage applies reinforcement learning to optimize the model's reasoning abilities through reward functions based on physical conditions. Our approach allows the model to dynamically adjust and improve the physical properties of generated videos, ensuring adherence to physical laws. Experimental results demonstrate that PhysAR can generate videos that are physically consistent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。