arXiv:2410.15957cs.CV2024-10被引 110

通过稀疏对极注意力提升图像到视频生成的相机控制精度

CamI2V: Camera-Controlled Image-to-Video Diffusion Model

  • 用对极线注意力仅聚合对应特征,降低跨帧噪声干扰
  • 在RealEstate10K上相机控制力提升25.64%,兼顾动态与画质
  • 适合需要精准相机控制的视频生成研究者

近期进展将相机位姿作为用户友好且符合物理规律的条件引入视频扩散模型,实现精确相机控制。本文识别出关键挑战在于有效建模噪声跨帧交互以增强几何一致性与相机可控性。我们创新性地将条件质量与不确定性降低关联,将噪声跨帧特征视为一种含噪条件。鉴于噪声条件虽提供确定性信息但亦引入随机性和误导风险,提出仅沿对应对极线聚合特征的对极注意力机制,以获取最优噪声条件量。针对对极线因快速相机运动、动态物体或遮挡消失的问题,设计鲁棒处理策略,确保在多样环境下的性能。此外,构建更稳健可复现的评估流程,解决现有相机控制指标的不准确与不稳定问题。方法在RealEstate10K数据集上实现25.64%的相机控制力提升,未牺牲动态表现或生成质量,并展现强域外泛化能力。训练与推理分别仅需24GB和12GB显存,支持16帧、256x256分辨率序列。所有检查点及训练评估代码将公开。

原文摘要 · Abstract (English)

Recent advancements have integrated camera pose as a user-friendly and physics-informed condition in video diffusion models, enabling precise camera control. In this paper, we identify one of the key challenges as effectively modeling noisy cross-frame interactions to enhance geometry consistency and camera controllability. We innovatively associate the quality of a condition with its ability to reduce uncertainty and interpret noisy cross-frame features as a form of noisy condition. Recognizing that noisy conditions provide deterministic information while also introducing randomness and potential misguidance due to added noise, we propose applying epipolar attention to only aggregate features along corresponding epipolar lines, thereby accessing an optimal amount of noisy conditions. Additionally, we address scenarios where epipolar lines disappear, commonly caused by rapid camera movements, dynamic objects, or occlusions, ensuring robust performance in diverse environments. Furthermore, we develop a more robust and reproducible evaluation pipeline to address the inaccuracies and instabilities of existing camera control metrics. Our method achieves a 25.64% improvement in camera controllability on the RealEstate10K dataset without compromising dynamics or generation quality and demonstrates strong generalization to out-of-domain images. Training and inference require only 24GB and 12GB of memory, respectively, for 16-frame sequences at 256x256 resolution. We will release all checkpoints, along with training and evaluation code. Dynamic videos are best viewed at https://zgctroy.github.io/CamI2V.

图像到视频相机控制扩散模型对极注意力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。