arXiv:2604.06010cs.CV2026-04

让视频生成同时控制画面内容和镜头运动,自由组合更灵活。

OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control

  • 分离画面内容与镜头运动,实现任意搭配生成
  • 在复杂镜头动作下仍保持高质量视觉效果
  • 适合需要精细镜头控制的视频创作人群

视频本质上融合了场景动态内容与观察视角的相机运动两个核心维度。然而现有生成模型常将两者混淆,限制独立控制。本文提出OmniCamera统一框架,显式解耦并操控这两个维度。通过组合式设计,支持任意内容与镜头条件的自由配对,实现前所未有的创作灵活性。为解决该系统固有的模态冲突与数据稀缺问题,提出两项关键创新:首先构建OmniCAM——一个结合精心筛选的真实视频与合成数据的混合数据集,提供多样化的成对样本以支撑多任务学习;其次提出双层课程协同训练策略,从条件难度和数据来源两方面逐步优化:先按控制难度渐进引入模态,再在合成数据上精准训练后迁移至真实数据以实现逼真视觉效果。实验表明,OmniCamera达到当前最优性能,在复杂相机运动下仍保持卓越视觉质量。

原文摘要 · Abstract (English)

Video fundamentally intertwines two crucial axes: the dynamic content of a scene and the camera motion through which it is observed. However, existing generation models often entangle these factors, limiting independent control. In this work, we introduce OmniCamera, a unified framework designed to explicitly disentangle and command these two dimensions. This compositional approach enables flexible video generation by allowing arbitrary pairings of camera and content conditions, unlocking unprecedented creative control. To overcome the fundamental challenges of modality conflict and data scarcity inherent in such a system, we present two key innovations. First, we construct OmniCAM, a novel hybrid dataset combining curated real-world videos with synthetic data that provides diverse paired examples for robust multi-task learning. Second, we propose a Dual-level Curriculum Co-Training strategy that mitigates modality interference and synergistically learns from diverse data sources. This strategy operates on two levels: first, it progressively introduces control modalities by difficulties (condition-level), and second, trains for precise control on synthetic data before adapting to real data for photorealism (data-level). As a result, OmniCamera achieves state-of-the-art performance, enabling flexible control for complex camera movements while maintaining superior visual quality.

视频生成相机控制多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。