arXiv:2409.01502cs.CVcs.AI2024-09被引 1

用3D角色驱动扩散模型,实现多人物视频精准控制与高真实感生成。

AMG: Avatar Motion Guided Video Generation

  • 通过3D角色渲染引导扩散模型生成视频,融合2D真实感与3D可控性。
  • 支持多人物、相机位姿、动作和背景风格的精确控制,生成视频更自然。
  • 首次实现基于动态摄像头视频重建角色动作并生成高质量视频,适合影视特效应用。

随着深度生成模型的发展,人类视频生成任务受到广泛关注。由于人体拓扑结构复杂且对视觉伪影敏感,生成逼真的人类运动视频具有天然挑战性。现有的2D媒体生成方法虽利用海量人类数据集,但难以实现3D感知控制;而基于3D角色的方法虽控制自由度高,却缺乏真实感,难以与背景场景无缝融合。本文提出AMG,通过将视频扩散模型条件于3D角色的受控渲染,结合2D真实感与3D可控性。我们还引入一种新型数据处理流程,从动态摄像头视频中重建并渲染人体动作。AMG是首个实现多人物扩散视频生成,并能精确控制相机位置、人体动作与背景风格的方法。大量评估表明,其在真实感和适应性方面优于基于姿态序列或驱动视频的现有方法。

原文摘要 · Abstract (English)

Human video generation task has gained significant attention with the advancement of deep generative models. Generating realistic videos with human movements is challenging in nature, due to the intricacies of human body topology and sensitivity to visual artifacts. The extensively studied 2D media generation methods take advantage of massive human media datasets, but struggle with 3D-aware control; whereas 3D avatar-based approaches, while offering more freedom in control, lack photorealism and cannot be harmonized seamlessly with background scene. We propose AMG, a method that combines the 2D photorealism and 3D controllability by conditioning video diffusion models on controlled rendering of 3D avatars. We additionally introduce a novel data processing pipeline that reconstructs and renders human avatar movements from dynamic camera videos. AMG is the first method that enables multi-person diffusion video generation with precise control over camera positions, human motions, and background style. We also demonstrate through extensive evaluation that it outperforms existing human video generation methods conditioned on pose sequences or driving videos in terms of realism and adaptability.

视频生成3D控制扩散模型多人物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。