arXiv:2505.21653cs.CV2025-05被引 8

让视频生成符合物理规律,用大模型推理规则并纠正错误

Think Before You Diffuse: Infusing Physical Rules into Video Diffusion

  • 用大语言模型从文本中推断物理规则,指导视频生成过程
  • 在扩散模型中间阶段验证物理合理性,提升生成结果真实度
  • 适合需要逼真物理效果的视频生成场景,如动画、仿真

近期视频扩散模型虽能生成视觉美观的视频,但在合成正确物理现象方面仍具挑战。真实世界运动、交互与动态的复杂性使从数据中学习物理规律十分困难。本文提出DiffPhy框架,通过微调预训练视频扩散模型,实现物理准确且照片级真实的视频生成。该方法利用大语言模型(LLM)从文本提示中推断丰富的物理上下文,并通过多模态大语言模型(MLLM)对中间潜在变量进行物理规则验证,引导梯度更新。将LLM的文本输出转化为连续信号,构建联合训练目标以兼顾物理准确性与文本语义对齐。同时通过注意力注入修正常见物理错误。我们还构建了一个高质量物理视频数据集,包含多样化的物理动作与事件,支持有效微调。大量实验表明,DiffPhy在多个公开基准上均达到领先性能。

原文摘要 · Abstract (English)

Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions, interactions, and dynamics introduce great difficulties when learning physics from data. In this work, we propose DiffPhy, a generic framework that enables physically-correct and photo-realistic video generation by fine-tuning a pre-trained video diffusion model. Our method leverages large language models (LLMs) to infer rich physical context from the text prompt. To incorporate this context into the video diffusion model, we use a multimodal large language model (MLLM) to verify intermediate latent variables against the inferred physical rules, guiding the gradient updates of model accordingly. Textual output of LLM is transformed into continuous signals. We then formulate a set of training objectives that jointly ensure physical accuracy and semantic alignment with the input text. Additionally, failure facts of physical phenomena are corrected via attention injection. We also establish a high-quality physical video dataset containing diverse phyiscal actions and events to facilitate effective finetuning. Extensive experiments on public benchmarks demonstrate that DiffPhy is able to produce state-of-the-art results across diverse physics-related scenarios. Our project page is available at https://bwgzk-keke.github.io/DiffPhy/.

视频生成扩散模型物理模拟大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。