用大模型引导视频生成,让动作更符合物理规律。
Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
- 通过多模态思维链迭代优化生成提示,融入物理常识
- 在PhyIQ基准上物理智商得分从56.31提升至62.38
- 无需训练、可适配多种模型,适合追求真实感的开发者
视频生成技术虽已取得显著视觉效果,但现有模型仍难以生成符合真实物理规律的结果。为此,我们提出一种基于大语言模型与视觉语言模型的迭代自精炼框架,利用物理一致性反馈指导视频生成。具体而言,引入多模态思维链(MM-CoT)过程,根据物理矛盾反馈逐步优化生成提示,从而提升结果质量。该方法无需训练、即插即用,可广泛适配各类视频生成模型。在PhyIQ基准上的实验表明,该方法将物理智商(Physics-IQ)得分从56.31提升至62.38。我们希望此项工作能为物理一致性视频生成提供初步探索,并为后续研究提供参考。
原文摘要 · Abstract (English)
Recent progress in video generation has led to impressive visual quality, yet current models still struggle to produce results that align with real-world physical principles. To this end, we propose an iterative self-refinement framework that leverages large language models and vision-language models to provide physics-aware guidance for video generation. Specifically, we introduce a multimodal chain-of-thought (MM-CoT) process that refines prompts based on feedback from physical inconsistencies, progressively enhancing generation quality. This method is training-free and plug-and-play, making it readily applicable to a wide range of video generation models. Experiments on the PhyIQ benchmark show that our method improves the Physics-IQ score from 56.31 to 62.38. We hope this work serves as a preliminary exploration of physics-consistent video generation and may offer insights for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。