arXiv:2412.01821cs.CV2024-12CVPR被引 62

用3D坐标显式建模,让视频生成更符合物理世界。

World-consistent Video Diffusion with Explicit 3D Modeling

  • 通过RGB与XYZ图像联合训练扩散模型,显式学习3D结构。
  • 单图转3D、多视角重建、相机控制生成等任务统一实现。
  • 支持灵活修复和新视角生成,适合3D内容创作场景。

扩散模型在图像与视频生成上取得突破,但在3D一致性生成方面仍存挑战。为此,我们提出世界一致的视频扩散模型(WVD),引入显式3D监督——利用XYZ图像为每个像素编码全局3D坐标。具体地,训练一个扩散Transformer以学习RGB与XYZ帧的联合分布。该方法通过灵活的修复策略支持多任务适配:例如,从真实RGB图像估计XYZ帧,或基于指定相机轨迹的XYZ投影生成新视角视频帧。此框架统一了单图转3D、多视图立体重建及相机控制视频生成等任务。在多个基准测试中表现优异,仅需一个预训练模型即可实现可扩展的3D一致图像与视频生成。

原文摘要 · Abstract (English)

Recent advancements in diffusion models have set new benchmarks in image and video generation, enabling realistic visual synthesis across single- and multi-frame contexts. However, these models still struggle with efficiently and explicitly generating 3D-consistent content. To address this, we propose World-consistent Video Diffusion (WVD), a novel framework that incorporates explicit 3D supervision using XYZ images, which encode global 3D coordinates for each image pixel. More specifically, we train a diffusion transformer to learn the joint distribution of RGB and XYZ frames. This approach supports multi-task adaptability via a flexible inpainting strategy. For example, WVD can estimate XYZ frames from ground-truth RGB or generate novel RGB frames using XYZ projections along a specified camera trajectory. In doing so, WVD unifies tasks like single-image-to-3D generation, multi-view stereo, and camera-controlled video generation. Our approach demonstrates competitive performance across multiple benchmarks, providing a scalable solution for 3D-consistent video and image generation with a single pretrained model.

视频生成3D建模扩散模型多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。