用文字或视频控制相机运动,生成高质量连贯视频
OmniCam: Unified Multimodal Video Generation via Camera Control
- 结合大模型与扩散模型,支持多模态输入控制相机轨迹
- 在多个指标上达到当前最佳效果,生成长序列视频保持时序一致
- 适合需要精细相机控制的影视创作与虚拟拍摄场景
相机控制通过改变相机位置和姿态实现多样视觉效果,受到广泛关注。然而现有方法存在交互复杂、控制能力有限等问题。为此,我们提出OmniCam统一多模态相机控制框架。利用大语言模型与视频扩散模型,OmniCam可生成时空一致的视频。支持多种输入组合:用户可提供文本或视频作为相机路径引导,以及图像或视频作为内容参考,实现对相机运动的精准控制。为支持OmniCam训练,我们构建了OmniTr数据集,包含大量高质量长序列轨迹、视频及对应描述。实验表明,该模型在多种指标下均达到当前最优水平,显著提升高质量相机控制视频生成性能。
原文摘要 · Abstract (English)
Camera control, which achieves diverse visual effects by changing camera position and pose, has attracted widespread attention. However, existing methods face challenges such as complex interaction and limited control capabilities. To address these issues, we present OmniCam, a unified multimodal camera control framework. Leveraging large language models and video diffusion models, OmniCam generates spatio-temporally consistent videos. It supports various combinations of input modalities: the user can provide text or video with expected trajectory as camera path guidance, and image or video as content reference, enabling precise control over camera motion. To facilitate the training of OmniCam, we introduce the OmniTr dataset, which contains a large collection of high-quality long-sequence trajectories, videos, and corresponding descriptions. Experimental results demonstrate that our model achieves state-of-the-art performance in high-quality camera-controlled video generation across various metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。