arXiv:2607.25321cs.AI2026-07

让视频生成更符合流体物理规律,通过模拟数据和双流光流监督提升运动真实性。

Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

论文配图:Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision
图 1 · 摘自论文原文
  • 用1638万条模拟流体视频+2320条真实视频构建新数据集,含测试集与文本生成基准。
  • 双流架构在预训练模型上新增光流解码器,仅更新解码器,提升运动一致性,最大提升8.75分。
  • 模型生成的流体运动更符合物理规律,人眼评估更偏好,光流误差低至0.54像素。

视频扩散模型虽视觉效果逼真,但在涉及流体时常违背基本物理规律:液柱空中断裂、倒水时水面不升、飞溅无动量与重力约束。我们归因于大规模视频-文本语料几乎无显式运动监督,模型仅学习流体外观而非动态。为此提出两项贡献:一,构建融合1638万条MPM模拟倾倒/晃动视频与2320条关键词筛选的真实倾倒视频的数据集,并包含1515条真实视频测试集及18个提示的文本到首帧泛化基准;二,提出基于预训练扩散-变压器视频生成器的双流图像到视频架构,增加轻量级光流解码分支,通过端点误差与平滑性损失显式训练,以零初始化卷积融合至RGB流,保持预训练主干不变。仅解码器更新,编码器、时序变换器与文本编码器冻结。在两个模型规模(1.3B与14B)及两个测试集上,相较冻结主干,视频物理常识与画质得分分别提升最高8.75与4.65分,优于领先开源对手,盲评中更受人类青睐。直接光流读出评估显示分布内端点误差低至0.54像素,证明模型已内化一致运动先验,非仅改善表面外观。

原文摘要 · Abstract (English)

Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liquid columns break apart in mid-air, container water levels fail to rise as liquid is poured in, and splashes disperse without regard to momentum or gravity. We attribute this gap to the fact that large-scale video-text corpora contain almost no explicit motion supervision, so models learn to imitate fluid appearance rather than dynamics. We address this with two contributions. First, we build a physics-simulation fluid dataset combining 1,638 MPM-simulated pouring/sloshing videos with 2,320 keyword-filtered real pouring videos mined from stock footage, plus two held-out test sets: a 1,515-video real-video benchmark and an 18-prompt text-to-first-frame generalization benchmark. Second, we introduce a dual-stream image-to-video architecture built on a pretrained diffusion-transformer video generator. It augments the standard RGB decoder with a lightweight Optical-Flow Decoder branch trained with explicit end-point-error and smoothness losses, fused into the RGB stream via zero-initialized convolutions so the pretrained backbone starts undisturbed. Only the two decoders are updated; the encoder, temporal transformer, and text encoder remain frozen. Across two model scales (1.3B and 14B) and two test sets, our method improves VideoPhy-2 Physical-Commonsense and Video-Quality scores over the frozen backbone by up to 8.75 and 4.65 points, outperforms a leading open competitor, and is preferred by human raters in a blind study. A direct optical-flow read-out evaluation further shows an end-point error as low as 0.54 pixels in-distribution, confirming the model has internalized a coherent motion prior rather than merely improving surface appearance.

视频生成流体模拟扩散模型物理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。