arXiv:2609.08215cs.CV2026-09

通过运动先验提升视频生成的物理合理性。

PhysFlow: Physics-Aware Optical Flow for Motion Controllable Video Generation

论文配图:PhysFlow: Physics-Aware Optical Flow for Motion Controllable Video Generation
图 1 · 摘自论文原文
  • 分两阶段生成:先建模物理驱动的光流,再据此合成外观。
  • 在10K物体、5万视频的物理数据集上训练,显著提升真实性。
  • 适合关注视频真实动态、物理模拟的研究者。

视频生成模型近年来受到广泛关注,因其能生成视觉吸引人的视频,但确保物理上一致且合理的运动仍是一大挑战,推动了物理真实性研究的兴起。为应对这一挑战,我们提出PhysFlow,一种新型两阶段框架,通过将视频生成分解为运动感知的光流生成和运动条件下的外观合成,以提升生成视频的物理合理性。具体而言,PhysFlow包括一个物理感知光流生成器PA-Flow和一个光流引导的视频生成器FlowRender。第一阶段中,PA-Flow采用物理感知注意力模块,分别建模运动属性对全局运动和材料属性对局部形变的影响,生成显式表示运动的光流视频。第二阶段,FlowRender利用解耦的运动表示作为指导,合成真实的纹理与外观,最终生成具有物理合理性的视频。为进一步支持带显式物理监督的模型训练,我们构建了PhysVideo,一个基于物理引擎和3D-GS渲染生成的物理视频数据集,包含10,000个前景物体和50,000条真实视频序列,并标注了运动与材料属性。大量实验表明,与现有方法相比,PhysFlow在保持高视觉保真度的同时,生成的视频具有更优的物理合理性。

原文摘要 · Abstract (English)

Video generation models have recently attracted substantial attention for their ability to generate visually compelling videos, yet ensuring physically consistent and plausible dynamics still remains a fundamental challenge, driving a growing line of research on physical realism in video generation. To address this challenge, motivated by the fact that physical regularities are primarily encoded in motion patterns, we propose PhysFlow, a novel two-stage framework for improving the physical plausibility of generated videos by decomposing video generation into motion-aware optical flow generation followed by motion-conditioned appearance synthesis. Specifically, PhysFlow consists of a physics-aware optical-flow video generator called PA-Flow and a flow-guided video generator called FlowRender. During the first stage, PA-Flow employs a physics-aware attention module to model how motion attributes and material properties influence global motion and local deformation, respectively, and generates an optical flow video as an explicit representation of motion. In the second stage, FlowRender leverages the decoupled motion representation as guidance to synthesize realistic textures and appearances, ultimately producing the final physically plausible video. To further support model training with explicit physical supervision, we construct PhysVideo, a physics-based video dataset generated with a physics engine and 3D-GS rendering, containing 10K foreground objects and 50K realistic video sequences with annotations of motion and material properties. Extensive experiments demonstrate that our proposed PhysFlow generates videos with superior physical plausibility while maintaining high visual fidelity compared with existing methods.

视频生成物理模拟光流

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。