arXiv:2503.23368cs.CVcs.AI2025-03ICCV被引 47

用视觉语言模型引导视频生成,让动画更符合物理规律。

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

  • 先用视觉语言模型规划粗略运动轨迹,融入物理常识推理
  • 再以轨迹为指导生成视频,加入噪声保留细节自由度
  • 生成动作更符合真实物理,适合需要合理动态的应用

视频扩散模型(VDM)近年来发展迅速,能生成高度逼真的视频,被视为潜在的世界模拟器。然而,由于缺乏对物理规律的理解,现有模型常产生不符合物理的动态和事件序列。为此,我们提出一种两阶段图像到视频生成框架,显式融合视觉与语言引导的物理先验。第一阶段使用视觉语言模型(VLM)作为粗粒度运动规划器,结合思维链与物理感知推理,预测近似真实物理动态的粗略运动轨迹,确保帧间一致性。第二阶段利用预测的运动轨迹指导VDM生成视频,因轨迹本身较粗糙,在推理时添加噪声,赋予VDM生成更精细运动的自由度。大量实验表明,该框架可生成更符合物理规律的运动,对比评估显示其显著优于现有方法。更多结果见项目页面:https://madaoer.github.io/projects/physically_plausible_video_generation。

原文摘要 · Abstract (English)

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos due to an inherent lack of understanding of physics, resulting in incorrect dynamics and event sequences. To address this limitation, we propose a novel two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior. In the first stage, we employ a Vision Language Model (VLM) as a coarse-grained motion planner, integrating chain-of-thought and physics-aware reasoning to predict a rough motion trajectories/changes that approximate real-world physical dynamics while ensuring the inter-frame consistency. In the second stage, we use the predicted motion trajectories/changes to guide the video generation of a VDM. As the predicted motion trajectories/changes are rough, noise is added during inference to provide freedom to the VDM in generating motion with more fine details. Extensive experimental results demonstrate that our framework can produce physically plausible motion, and comparative evaluations highlight the notable superiority of our approach over existing methods. More video results are available on our Project Page: https://madaoer.github.io/projects/physically_plausible_video_generation.

视频生成物理模拟视觉语言模型扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。