arXiv:2504.21855cs.CV2025-04被引 4

用3D运动先验提升视频生成质量,让动作更真实连贯。

ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling

  • 引入参数化3D运动先验,显式建模物体运动轨迹。
  • 仅1.5B参数即超越13B参数的先进模型,复杂动作生成更精准。
  • 适配现有扩散模型,无需重训练,适合追求高效高质视频生成者。

近年来视频生成取得显著进展,但在复杂动作与交互生成方面仍面临挑战。为此,我们提出ReVision,一个即插即用框架,将参数化3D模型知识显式融入预训练条件视频生成模型,显著提升生成高质量、复杂运动与交互视频的能力。ReVision包含三个阶段:首先,使用视频扩散模型生成粗略视频;其次,从粗略视频中提取2D和3D特征,构建3D物体中心表示,并通过提出的参数化运动先验模型优化,生成精确的3D运动序列;最后,将优化后的运动序列作为额外条件反馈至原视频扩散模型,生成运动一致的视频,即使在复杂动作与交互场景下也能保持一致性。我们在Stable Video Diffusion上验证了该方法的有效性,ReVision显著提升了运动保真度与连贯性。值得注意的是,仅1.5B参数的ReVision在复杂视频生成任务中,远超参数量超过13B的先进模型。结果表明,通过融合3D运动知识,即使是小型视频扩散模型也能生成更具真实感与可控性的复杂动作与交互,为物理合理视频生成提供了一种有前景的解决方案。

原文摘要 · Abstract (English)

In recent years, video generation has seen significant advancements. However, challenges still persist in generating complex motions and interactions. To address these challenges, we introduce ReVision, a plug-and-play framework that explicitly integrates parameterized 3D model knowledge into a pretrained conditional video generation model, significantly enhancing its ability to generate high-quality videos with complex motion and interactions. Specifically, ReVision consists of three stages. First, a video diffusion model is used to generate a coarse video. Next, we extract a set of 2D and 3D features from the coarse video to construct a 3D object-centric representation, which is then refined by our proposed parameterized motion prior model to produce an accurate 3D motion sequence. Finally, this refined motion sequence is fed back into the same video diffusion model as additional conditioning, enabling the generation of motion-consistent videos, even in scenarios involving complex actions and interactions. We validate the effectiveness of our approach on Stable Video Diffusion, where ReVision significantly improves motion fidelity and coherence. Remarkably, with only 1.5B parameters, it even outperforms a state-of-the-art video generation model with over 13B parameters on complex video generation by a substantial margin. Our results suggest that, by incorporating 3D motion knowledge, even a relatively small video diffusion model can generate complex motions and interactions with greater realism and controllability, offering a promising solution for physically plausible video generation.

视频生成扩散模型3D运动可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。