让视频生成更懂场景,还能精准控制镜头运动
CamC2V: Context-aware Controllable Video Generation
- 用多图+3D约束+镜头控制,提升视频上下文理解
- 在RealEstate10K上FVD降低24.09%,画质与镜头可控性双提升
- 适合需要真实场景还原的影视生成与虚拟拍摄
近期的图像到视频(I2V)扩散模型在场景理解与生成质量上表现优异,能利用图像条件引导生成。然而,这些模型主要仅对静态图像进行动画化,难以扩展至原始提供的上下文之外。引入额外约束如相机轨迹虽可增强多样性,但常导致视觉质量下降,限制了对忠实场景再现有要求的任务应用。本文提出CamC2V,一种将多张图像作为上下文,并结合3D约束与相机控制的上下文到视频(C2V)模型,以丰富全局语义与细粒度视觉细节,实现更连贯、上下文感知的视频生成。此外,我们强调时间感知对于有效上下文表征的重要性。在RealEstate10K数据集上的综合实验表明,该方法在视觉质量上相较基线提升24.09%(FVD),同时增强相机可控性。代码已公开于https://github.com/LDenninger/CamC2V。
原文摘要 · Abstract (English)
Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However, these models primarily animate static images without extending beyond their provided context. Introducing additional constraints, such as camera trajectories, can enhance diversity but often degrade visual quality, limiting their applicability for tasks requiring faithful scene representation. We propose CamC2V, a context-to-video (C2V) model that integrates multiple image conditions as context with 3D constraints alongside camera control to enrich both global semantics and fine-grained visual details. This enables more coherent and context-aware video generation. Moreover, we motivate the necessity of temporal awareness for an effective context representation. Our comprehensive study on the RealEstate10K dataset demonstrates a $24.09\%$ (FVD) improvement in visual quality and camera controllability. Our code is publicly available at: https://github.com/LDenninger/CamC2V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。