无需估计相机位姿,用可见区域掩码生成精准锚视频,实现高效视频相机控制。
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
- 基于首帧可见性掩码生成对齐锚视频,免去点云与相机轨迹估计。
- 训练参数减少99%以上,仅需少量数据和步骤即达当前最佳性能。
- 支持零样本跨视频场景迁移,适合真实世界视频的相机控制应用。
现有视频生成中相机控制方法常通过渲染估计点云的轨迹生成锚视频作为扩散模型的结构化先验,但点云与相机轨迹估计误差导致锚视频不准,增加训练成本与效率低下。为此,本文提出EPiC框架,无需相机位姿或点云估计即可构建精确对齐的锚视频。具体而言,通过首帧可见性掩码生成高精度锚视频,确保强对齐并消除估计依赖,适用于任意真实视频。同时引入轻量级Anchor-ControlNet模块,仅增加少于1%参数,将锚视频引导融入预训练视频扩散模型。EPiC实现高效训练,显著降低参数、训练步数与数据需求,并在测试时对点云生成的锚视频具备鲁棒泛化能力,实现精准3D感知相机控制。在RealEstate10K与MiraData数据集上达到当前最优表现,且展现强大零样本视频到视频迁移能力。
原文摘要 · Abstract (English)
Recent approaches for video generation with camera control often create anchor videos (i.e., rendered videos that approximate desired camera motions) to guide diffusion models as a structured prior, by rendering from estimated point clouds following camera trajectories. However, errors in point cloud and camera trajectory estimation often lead to inaccurate anchor videos with higher training cost and low efficiency, as the model is forced to compensate for rendering misalignments. To address these limitations, we introduce EPiC, an efficient and precise camera control learning framework that constructs well-aligned training anchor videos without the need for camera pose or point cloud estimation. Concretely, we create highly precise anchor videos by masking source videos based on first-frame visibility, which ensures strong alignment, eliminates the need for camera/point cloud estimation, and thus can be readily applied to any in-the-wild video. Furthermore, we introduce Anchor-ControlNet, a lightweight module that integrates anchor video guidance in visible regions to pretrained video diffusion models, with less than 1% of additional parameters. EPiC achieves efficient training with substantially fewer parameters, training steps, and less data, and generalizes robustly to anchor videos made with point clouds at test time, enabling precise 3D-informed camera control. EPiC achieves SoTA performance on RealEstate10K and MiraData for I2V camera control task. Notably, EPiC also exhibits strong zero-shot generalization to video-to-video (V2V) scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。