让视频生成更符合物理规律,用真实视频引导学习。
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
- 用视觉语言模型构建大规模物理增强数据集。
- 基于真实视频优化,显著提升物理一致性表现。
- 轻量级设计支持高效训练,适合研究与应用。
近期文本到视频(T2V)生成在视觉质量上取得进展,但忠实遵循物理规律仍是挑战。现有方法依赖图形或提示扩展,在复杂环境泛化能力差,且缺乏丰富物理交互的训练数据。本文提出物理增强数据构建管道PhyAugPipe,利用具有思维链推理能力的视觉语言模型(VLM)收集大规模数据集PhyVidGen-135K。进一步设计物理感知分组直接偏好优化框架PhyGDPO,以真实视频为优胜案例确保正确物理学习,并采用分组Plackett-Luce概率模型捕捉整体偏好。引入基于VLM的物理奖励机制(PGR),聚焦困难物理场景;提出LoRA-Switch参考(LoRA-SR),避免全模型复制,实现高效DPO训练。实验表明,该方法在PhyGenBench和VideoPhy2基准上显著优于现有开源方法。
原文摘要 · Abstract (English)
Recent advances in text-to-video (T2V) generation have achieved good visual quality, yet synthesizing videos that faithfully follow physical laws remains an open challenge. Existing methods mainly based on graphics or prompt extension struggle to generalize beyond simple simulated environments or learn implicit physical reasoning. The scarcity of training data with rich physics interactions and phenomena is also a problem. In this paper, we first introduce a Physics-Augmented video data construction Pipeline, PhyAugPipe, that leverages a vision-language model (VLM) with chain-of-thought reasoning to collect a large-scale training dataset, PhyVidGen-135K. Then we formulate a principled Physics-aware Groupwise Direct Preference Optimization, PhyGDPO, framework that uses real-world video as winning case to guarantee correct physics learning and builds upon the groupwise Plackett-Luce probabilistic model to capture holistic preferences beyond pairwise comparisons. In PhyGDPO, we design a Physics-Guided Rewarding (PGR) scheme that leverages VLM-based physical rewards to direct the optimization to focus on challenging physics cases. In addition, we propose a LoRA-Switch Reference (LoRA-SR) scheme that avoids full-model duplication as reference for efficient DPO training. Experiments show that our method significantly outperforms state-of-the-art open-source methods on PhyGenBench and VideoPhy2. Please check our project page at https://caiyuanhao1998.github.io/project/PhyGDPO for more video results. Our code, data, and models are publicly available at https://github.com/caiyuanhao1998/Open-PhyGDPO
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。