让视频生成能精准理解用户意图的相机运动
CT-1: Vision-Language-Camera Models Transfer Spatial Reasoning Knowledge to Camera-Controllable Video Generation
- 用视觉语言模型+扩散变换器,从文本推断真实相机轨迹
- 在4700万帧数据上训练,相机控制准确率提升25.7%
- 适合做自动化视频创作或虚拟拍摄的开发者
相机可控视频生成旨在合成具有灵活且物理合理的相机运动的视频。然而,现有方法要么仅通过文本提示实现粗糙的相机控制,要么依赖耗时的手动相机轨迹参数,限制了其在自动化场景中的应用。为此,我们提出一种新型视觉-语言-相机模型CT-1(Camera Transformer 1),该模型通过精确估计相机轨迹,将空间推理知识迁移至视频生成。基于视觉-语言模块与扩散变换器构建,CT-1在频域中采用小波正则化损失,有效学习复杂相机轨迹分布。这些轨迹被整合进视频扩散模型,实现与用户意图一致的空间感知相机控制。为支持CT-1训练,我们设计专用数据整理流程,构建了包含超过4700万帧的大规模数据集CT-200K。实验表明,该框架成功弥合了空间推理与视频合成之间的差距,生成的视频在相机控制精度上相比之前方法提升25.7%,具备高保真度与真实性。
原文摘要 · Abstract (English)
Camera-controllable video generation aims to synthesize videos with flexible and physically plausible camera movements. However, existing methods either provide imprecise camera control from text prompts or rely on labor-intensive manual camera trajectory parameters, limiting their use in automated scenarios. To address these issues, we propose a novel Vision-Language-Camera model, termed CT-1 (Camera Transformer 1), a specialized model designed to transfer spatial reasoning knowledge to video generation by accurately estimating camera trajectories. Built upon vision-language modules and a Diffusion Transformer model, CT-1 employs a Wavelet-based Regularization Loss in the frequency domain to effectively learn complex camera trajectory distributions. These trajectories are integrated into a video diffusion model to enable spatially aware camera control that aligns with user intentions. To facilitate the training of CT-1, we design a dedicated data curation pipeline and construct CT-200K, a large-scale dataset containing over 47M frames. Experimental results demonstrate that our framework successfully bridges the gap between spatial reasoning and video synthesis, yielding faithful and high-quality camera-controllable videos and improving camera control accuracy by 25.7% over prior methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。