arXiv:2501.03847cs.CVcs.AI2025-01International Conf…被引 185

用3D追踪视频控制视频生成,实现相机、动作等多任务灵活调控。

Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control

论文配图:Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
图 1 · 摘自论文原文
  • 以3D追踪视频为输入,实现统一架构下的多类型视频控制。
  • 仅用10,000张视频微调3天,8张H800 GPU即可完成训练。
  • 支持网格转视频、相机控制、动作迁移等任务,生成更连贯视频。

扩散模型在文本或图像生成高质量视频方面表现优异,但对视频生成过程的精确控制(如相机运动或内容编辑)仍具挑战。现有方法通常仅支持单一控制类型,缺乏灵活性。本文提出扩散作为着色器(Diffusion as Shader, DaS),在统一架构中支持多种视频控制任务。核心思想是:视频本质上是动态3D内容的2D渲染,因此需依赖3D控制信号实现多样化控制。与仅使用2D信号的方法不同,DaS采用3D追踪视频作为控制输入,使生成过程天然具备3D感知能力。通过操控3D追踪视频,可实现多样化的视频控制。此外,3D追踪视频能有效关联帧间信息,显著提升生成视频的时间一致性。仅需在8张H800 GPU上进行3天微调,使用少于10,000张视频,DaS即在网格转视频、相机控制、动作迁移和物体操作等任务上展现出强大控制能力。

原文摘要 · Abstract (English)

Diffusion models have demonstrated impressive performance in generating high-quality videos from text prompts or images. However, precise control over the video generation process, such as camera manipulation or content editing, remains a significant challenge. Existing methods for controlled video generation are typically limited to a single control type, lacking the flexibility to handle diverse control demands. In this paper, we introduce Diffusion as Shader (DaS), a novel approach that supports multiple video control tasks within a unified architecture. Our key insight is that achieving versatile video control necessitates leveraging 3D control signals, as videos are fundamentally 2D renderings of dynamic 3D content. Unlike prior methods limited to 2D control signals, DaS leverages 3D tracking videos as control inputs, making the video diffusion process inherently 3D-aware. This innovation allows DaS to achieve a wide range of video controls by simply manipulating the 3D tracking videos. A further advantage of using 3D tracking videos is their ability to effectively link frames, significantly enhancing the temporal consistency of the generated videos. With just 3 days of fine-tuning on 8 H800 GPUs using less than 10k videos, DaS demonstrates strong control capabilities across diverse tasks, including mesh-to-video generation, camera control, motion transfer, and object manipulation.

视频生成扩散模型3D控制动作迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。