arXiv:2606.13432cs.CVcs.AI2026-06被引 5

无需配对数据,用网格视频实现多镜头相机动作克隆

OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data

论文配图:OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
图 1 · 摘自论文原文
  • 将相机参数转为网格运动视频,支持多镜头生成
  • 基于百万级数据训练,实现角色、动作、相机协同控制
  • 通过分层提示扩展机制,融合多种控制信号

从参考视频中克隆相机运动是视频生成中的关键任务,能提供直观精准的控制。现有方法或依赖参数化表示难以处理多镜头生成,或需合成跨配对数据,受限于数据稀缺导致复杂相机运动克隆性能不佳。为此,我们提出一种通用相机运动表示,将相机编码为网格运动视频,可视化呈现相机参数并支持多种轨迹集成,实现多镜头视频生成。基于此,我们构建OmniDirector框架,该框架在百万规模相机网格-视频对上训练,协调角色、动作与相机,为多模态扩散变换器提供导演级控制。此外,我们设计了新型分层提示扩展代理,通过系统理解信号间关系,协调描述相机运动与视觉内容。大量实验表明,本框架性能优异且可控性强。

原文摘要 · Abstract (English)

Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametric representations that fail to handle multi-shot generation or synthesize cross-paired data, which suffer from data scarcity, resulting in poor performance in complicated camera motion cloning. To address these issues, we introduce a general camera motion representation that encodes cameras as grid motion videos. This camera grid represents the camera parameters visually and supports the integration of diverse trajectories for multi-shot video generation. Building upon this, we propose OmniDirector, a unified framework trained on a million-scale camera grid-video pairs that coordinates characters, actions, and cameras to provide director-level control for multimodal diffusion transformers. Furthermore, we design a novel hierarchical prompt expansion agent that harmoniously integrates different control signals by systematically describing camera motion and visual content through understanding signal relationships. Extensive experiments demonstrate the superior performance and outstanding controllability of our framework. Project page: https://ymlinfeng.github.io/OmniDirector.github.io/

视频生成相机控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。