无需深度估计,通过无限单应性实现高精度相机控制视频生成
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
- 利用无限单应性将3D相机旋转编码到2D潜在空间
- 在合成数据上训练后可直接迁移到真实场景,相机姿态误差降低40%
- 适合需要精准镜头控制的影视后期与虚拟拍摄场景
视频扩散模型的发展推动了动态场景中相机控制的新视角视频生成研究,旨在为创作者提供后期制作中的电影级镜头控制能力。当前方法面临两大挑战:基于重投影的方法易受深度估计误差影响;现有数据集相机轨迹多样性不足,限制模型泛化能力。为此,本文提出InfCam框架,无需深度信息即可实现高姿态保真度的相机控制视频生成。该框架包含两个核心组件:(1) 无限单应性扭曲,在视频扩散模型的2D潜在空间中直接编码3D相机旋转,通过端到端训练预测残差视差项,确保姿态精度;(2) 数据增强流水线,将现有合成多视角数据转化为具有多样化轨迹和焦距的序列。实验表明,InfCam在相机姿态准确性和视觉保真度上优于基线方法,且能从合成数据有效泛化至真实场景。
原文摘要 · Abstract (English)
Recent progress in video diffusion models has spurred growing interest in camera-controlled novel-view video generation for dynamic scenes, aiming to provide creators with cinematic camera control capabilities in post-production. A key challenge in camera-controlled video generation is ensuring fidelity to the specified camera pose, while maintaining view consistency and reasoning about occluded geometry from limited observations. To address this, existing methods either train trajectory-conditioned video generation model on trajectory-video pair dataset, or estimate depth from the input video to reproject it along a target trajectory and generate the unprojected regions. Nevertheless, existing methods struggle to generate camera-pose-faithful, high-quality videos for two main reasons: (1) reprojection-based approaches are highly susceptible to errors caused by inaccurate depth estimation; and (2) the limited diversity of camera trajectories in existing datasets restricts learned models. To address these limitations, we present InfCam, a depth-free, camera-controlled video-to-video generation framework with high pose fidelity. The framework integrates two key components: (1) infinite homography warping, which encodes 3D camera rotations directly within the 2D latent space of a video diffusion model. Conditioning on this noise-free rotational information, the residual parallax term is predicted through end-to-end training to achieve high camera-pose fidelity; and (2) a data augmentation pipeline that transforms existing synthetic multiview datasets into sequences with diverse trajectories and focal lengths. Experimental results demonstrate that InfCam outperforms baseline methods in camera-pose accuracy and visual fidelity, generalizing well from synthetic to real-world data. Link to our project page:https://emjay73.github.io/InfCam/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。