提升视频生成的3D一致性,让物体在视角变化时保持稳定。
View-Consistent Diffusion Representations for 3D-Consistent Video Generation
- 通过学习多视角一致的扩散表示来增强视频生成的3D稳定性。
- 在多类视频生成任务中显著提升3D一致性,减少视角变化下的形变问题。
- 适合关注视频真实感与三维结构保真的研究者与开发者。
视频生成模型在生成逼真内容方面取得了显著进展,广泛应用于模拟、游戏和影视制作。然而,当前生成的视频仍存在因3D不一致导致的视觉伪影,例如物体在相机视角变化时发生形变,影响用户体验与模拟精度。受扩散模型表征对齐研究的启发,我们提出假设:提升视频扩散表征的多视角一致性可增强视频的3D一致性。通过对多个近期相机控制的视频扩散模型进行分析,我们发现3D一致性表征与视频生成质量之间存在强相关性。为此,我们提出ViCoDR方法,通过学习多视角一致的扩散表示来改善视频模型的3D一致性。我们在相机控制的图像到视频、文本到视频及多视角生成模型上评估ViCoDR,结果表明生成视频的3D一致性得到显著提升。项目页面:https://danier97.github.io/ViCoDR。
原文摘要 · Abstract (English)
Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D inconsistencies, e.g., objects and structures deforming under changes in camera pose, which can undermine user experience and simulation fidelity. Motivated by recent findings on representation alignment for diffusion models, we hypothesize that improving the multi-view consistency of video diffusion representations will yield more 3D-consistent video generation. Through detailed analysis on multiple recent camera-controlled video diffusion models we reveal strong correlations between 3D-consistent representations and videos. We also propose ViCoDR, a new approach for improving the 3D consistency of video models by learning multi-view consistent diffusion representations. We evaluate ViCoDR on camera controlled image-to-video, text-to-video, and multi-view generation models, demonstrating significant improvements in the 3D consistency of the generated videos. Project page: https://danier97.github.io/ViCoDR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。