无需相机标注,用统一符号实现精准视频摄像机控制。
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
- 用几何基础模型生成深度与相机参数,转为统一符号注入扩散模型。
- 分两阶段训练,实现几何一致、时间稳定且视觉真实的视频生成。
- 可扩展至图像修复和布局控制,适合需要灵活控制的视频生成场景。
视频生成中精确控制摄像机仍具挑战,现有方法依赖难以扩展且与深度估计不一致的相机位姿标注,导致训练-测试差异。我们提出CETCAM框架,通过一致且可扩展的标记化方案,消除对相机标注的需求。CETCAM利用VGGT等几何基础模型估计深度和相机参数,并将其转化为统一的、具有几何感知能力的标记。这些标记通过轻量级上下文模块无缝集成到预训练视频扩散模型中。模型采用两阶段训练:首先从多样化的原始视频数据中学习鲁棒的摄像机可控性,随后在高质量数据集上优化细节视觉质量。多基准测试结果表明,CETCAM在几何一致性、时间稳定性与视觉真实感方面均达到当前最优水平。此外,其对图像修复、布局控制等额外控制模态也表现出强适应性,凸显其超越摄像机控制的灵活性。项目页面见 https://sjtuytc.github.io/CETCam_project_page.github.io/。
原文摘要 · Abstract (English)
Achieving precise camera control in video generation remains challenging, as existing methods often rely on camera pose annotations that are difficult to scale to large and dynamic datasets and are frequently inconsistent with depth estimation, leading to train-test discrepancies. We introduce CETCAM, a camera-controllable video generation framework that eliminates the need for camera annotations through a consistent and extensible tokenization scheme. CETCAM leverages recent advances in geometry foundation models, such as VGGT, to estimate depth and camera parameters and converts them into unified, geometry-aware tokens. These tokens are seamlessly integrated into a pretrained video diffusion backbone via lightweight context blocks. Trained in two progressive stages, CETCAM first learns robust camera controllability from diverse raw video data and then refines fine-grained visual quality using curated high-fidelity datasets. Extensive experiments across multiple benchmarks demonstrate state-of-the-art geometric consistency, temporal stability, and visual realism. Moreover, CETCAM exhibits strong adaptability to additional control modalities, including inpainting and layout control, highlighting its flexibility beyond camera control. The project page is available at https://sjtuytc.github.io/CETCam_project_page.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。