用隐式科学知识引导扩散模型,从单帧生成符合物理规律的科学视频。
Latent Knowledge-Guided Video Diffusion for Scientific Phenomena Generation from a Single Initial Frame
- 通过自编码器和光流模型提取静态与动态科学知识
- 生成伪语言提示,提升视频生成的物理一致性
- 在流体模拟与台风观测中效果优于现有方法
视频扩散模型在自然场景生成中表现优异,但在流体模拟、气象过程等受科学定律支配的科学现象生成中表现不佳,主要因领域差异大、训练数据少且缺乏描述性标注。为此,本文提取隐式科学知识,提出新框架,使视频扩散模型能从单个初始帧生成科学现象。静态知识通过预训练掩码自编码器获取,动态知识则来自预训练光流预测模型。基于CLIP视觉与语言编码器对齐的空间关系,将受知识引导的视觉嵌入投影至空间与频域的伪语言提示嵌入。结合这些提示微调视频扩散模型,实现更符合科学规律的视频生成。在计算流体动力学模拟与真实台风观测数据上的实验表明,该方法在多种科学场景下均显著提升生成视频的保真度与一致性。
原文摘要 · Abstract (English)
Video diffusion models have achieved impressive results in natural scene generation, yet they struggle to generalize to scientific phenomena such as fluid simulations and meteorological processes, where underlying dynamics are governed by scientific laws. These tasks pose unique challenges, including severe domain gaps, limited training data, and the lack of descriptive language annotations. To handle this dilemma, we extracted the latent scientific phenomena knowledge and further proposed a fresh framework that teaches video diffusion models to generate scientific phenomena from a single initial frame. Particularly, static knowledge is extracted via pre-trained masked autoencoders, while dynamic knowledge is derived from pre-trained optical flow prediction. Subsequently, based on the aligned spatial relations between the CLIP vision and language encoders, the visual embeddings of scientific phenomena, guided by latent scientific phenomena knowledge, are projected to generate the pseudo-language prompt embeddings in both spatial and frequency domains. By incorporating these prompts and fine-tuning the video diffusion model, we enable the generation of videos that better adhere to scientific laws. Extensive experiments on both computational fluid dynamics simulations and real-world typhoon observations demonstrate the effectiveness of our approach, achieving superior fidelity and consistency across diverse scientific scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。