arXiv:2605.16795cs.CVcs.AI2026-05

用单图生成符合物理规律的视频,靠3D重建与仿真引导。

3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation

论文配图:3DPhysVideo: Consistency-Guided Flow SDE for Video Generation via 3D Scene Reconstruction and Physical Simulation
图 1 · 摘自论文原文
  • 用视频模型先重建360度3D场景,再用物理仿真引导生成视频
  • 在多物体和流体交互场景中生成高质量视频,性能超越现有方法
  • 无需训练,单张消费级显卡即可运行,适合快速生成物理真实视频

视频生成模型虽取得显著进展,但常产生违背物理动态的视觉伪影。现有工作如PhysGen3D通过网格重建和基于物理的渲染实现单图到3D物理建模,但在流体动力学、多物体交互和逼真度方面仍存挑战。本文提出3DPhysVideo,一种无需训练的新型流水线,可从单图生成物理真实的视频。我们复用现成的图像到视频(I2V)流模型分两阶段进行:首先,利用渲染点云引导I2V流模型作为新视角合成器,重建完整的360度3D场景几何;其次,在该几何上应用物理求解器后,将物理模拟的点云作为条件输入,再次引导同一I2V流模型生成最终高质量视频。提出的一致性引导流SDE将预测速度分解为去噪项与一致性偏差项,强制条件输入的一致性,从而有效复用模型完成3D重建与仿真引导视频生成。在包含多物体与流体交互的多样化实验中,本方法成功实现从单图到物理合理视频的跨越,且仅需单张消费级GPU即可高效运行。在GPT-based评分、VideoPhy基准测试及人工评估中均优于现有最先进基线。

原文摘要 · Abstract (English)

Video generative models have made remarkable progress, yet they often yield visual artifacts that violate grounding in physical dynamics. Recent works such as PhysGen3D tackle single image-to-3D physics through mesh reconstruction and Physically-Based Rendering, but challenges remain in modeling fluid dynamics, multi-object interactions and photorealism. This work introduces 3DPhysVideo, a novel training-free pipeline that generates physically realistic videos from a single image. We repurpose an off-the-shelf video model for two stages. First, we use it as a novel view synthesizer to reconstruct complete 360-degree 3D scene geometry by guiding the image-to-video (I2V) flow model with rendered point clouds. Second, after applying physics solvers to this geometry, the physically simulated point cloud is used to guide the same I2V flow model to synthesize final, high-quality videos. Consistency-Guided Flow SDE, which decomposes the predicted velocity of the I2V flow model into denoising and consistency bias, enforces consistency to the conditional inputs, allowing us to effectively repurpose the model for both 3D reconstruction and simulation-guided video generation. In the diverse experiments including multi-objects, and fluid interaction scenes, our method successfully bridges the gap from single-images to physically plausible videos, while remaining efficient to run on a single consumer GPU. It outperforms state-of-the-art baselines on GPT-based scores, VideoPhy benchmark and human evaluation.

视频生成物理仿真3D重建扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。