arXiv:2412.07760cs.CV2024-12被引 95

让多视角视频生成保持动态一致,支持任意相机位姿的虚拟拍摄。

SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

  • 在预训练文生视频模型上加同步模块,实现多视角内容一致性
  • 用混合数据训练,解决高质量多视角数据稀缺问题
  • 可重渲染新视角,适合虚拟影视与3D重建应用

视频扩散模型在模拟真实世界动态和保持三维一致性方面表现出色。受此启发,我们探索其在多视角间维持动态一致性的潜力,这在虚拟拍摄等场景中极具价值。不同于现有方法仅聚焦于单个物体的4D重建,本工作致力于从任意视角生成开放世界视频,并支持6自由度相机位姿。为此,提出一个即插即用模块,增强预训练文生视频模型以实现多相机视频生成,通过多视角同步模块保持外观与几何一致性。针对高质量训练数据稀缺问题,设计混合训练方案,融合多相机图像与单目视频来补充虚幻引擎渲染的多视角视频数据。此外,方法支持重渲染新视角等扩展功能。同时发布多视角同步视频数据集SynCamVideo-Dataset。

原文摘要 · Abstract (English)

Recent advancements in video diffusion models have shown exceptional abilities in simulating real-world dynamics and maintaining 3D consistency. This progress inspires us to investigate the potential of these models to ensure dynamic consistency across various viewpoints, a highly desirable feature for applications such as virtual filming. Unlike existing methods focused on multi-view generation of single objects for 4D reconstruction, our interest lies in generating open-world videos from arbitrary viewpoints, incorporating 6 DoF camera poses. To achieve this, we propose a plug-and-play module that enhances a pre-trained text-to-video model for multi-camera video generation, ensuring consistent content across different viewpoints. Specifically, we introduce a multi-view synchronization module to maintain appearance and geometry consistency across these viewpoints. Given the scarcity of high-quality training data, we design a hybrid training scheme that leverages multi-camera images and monocular videos to supplement Unreal Engine-rendered multi-camera videos. Furthermore, our method enables intriguing extensions, such as re-rendering a video from novel viewpoints. We also release a multi-view synchronized video dataset, named SynCamVideo-Dataset. Project page: https://jianhongbai.github.io/SynCamMaster/.

视频生成多视角扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。