让视频生成模型懂物理,提升真实感。
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation

- 将物理原理拆解为文本描述、定性类别和定量属性,注入生成过程。
- 在VideoPhy基准上显著提升物理合理性,优于现有模型。
- 适合需要高物理真实性的视频生成研究者使用。
近期文本到视频(T2V)生成技术如SoRA和Kling展现出构建世界模拟器的巨大潜力,但当前模型难以理解抽象物理规律,生成结果常违背物理定律。主要原因是物理原理与生成模型之间缺乏明确的指导信息。为此,我们提出世界模拟助手(WISA),一个将物理原理分解并融入T2V模型的有效框架。WISA将物理原理细分为文本描述、定性物理类别和定量物理属性,并引入混合物理专家注意力(MoPA)和物理分类器,增强模型的物理感知能力。此外,现有数据集中的物理现象多为弱表达或与其他过程纠缠,不适宜学习明确的物理规律。因此,我们构建了新型视频数据集WISA-32K,包含32,000个视频,覆盖动力学、热力学和光学三个领域的17条物理定律。实验表明,WISA能有效提升T2V模型与真实物理规律的一致性,在VideoPhy基准上实现显著改进。WISA与WISA-32K的可视化展示见https://360cvgroup.github.io/WISA/。
原文摘要 · Abstract (English)
Recent rapid advancements in text-to-video (T2V) generation, such as SoRA and Kling, have shown great potential for building world simulators. However, current T2V models struggle to grasp abstract physical principles and generate videos that adhere to physical laws. This challenge arises primarily from a lack of clear guidance on physical information due to a significant gap between abstract physical principles and generation models. To this end, we introduce the World Simulator Assistant (WISA), an effective framework for decomposing and incorporating physical principles into T2V models. Specifically, WISA decomposes physical principles into textual physical descriptions, qualitative physical categories, and quantitative physical properties. To effectively embed these physical attributes into the generation process, WISA incorporates several key designs, including Mixture-of-Physical-Experts Attention (MoPA) and a Physical Classifier, enhancing the model's physics awareness. Furthermore, most existing datasets feature videos where physical phenomena are either weakly represented or entangled with multiple co-occurring processes, limiting their suitability as dedicated resources for learning explicit physical principles. We propose a novel video dataset, WISA-32K, collected based on qualitative physical categories. It consists of 32,000 videos, representing 17 physical laws across three domains of physics: dynamics, thermodynamics, and optics. Experimental results demonstrate that WISA can effectively enhance the compatibility of T2V models with real-world physical laws, achieving a considerable improvement on the VideoPhy benchmark. The visual exhibitions of WISA and WISA-32K are available in the https://360cvgroup.github.io/WISA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。