arXiv:2411.17765cs.CV2024-11ICCV被引 20

统一控制视频生成中的多种运动指令,避免逻辑冲突

I2VControl: Disentangled and Unified Video Motion Synthesis Control

  • 将相机、物体拖拽、运动笔刷统一为点轨迹表示
  • 通过空间分区实现多控制类型动态协同,无冲突合成
  • 适配预训练模型,支持用户自由组合创意控制

运动可控性在视频生成中至关重要。然而,以往方法多局限于单一控制类型,组合时易产生逻辑冲突。本文提出 I2VControl 框架,通过重新定义相机控制、物体拖拽和运动笔刷,将其统一为基于点轨迹的一致表示,并设计空间分区策略,为每个单元分配对应控制类别,实现在单一生成流程中动态协调多种控制类型而无冲突。此外,我们设计了适配器结构,可插件式接入预训练模型且不依赖具体架构。大量实验表明,该方法在多种控制任务上表现优异,并支持用户驱动的创造性组合,显著提升创作自由度。项目页面:https://wanquanf.github.io/I2VControl

原文摘要 · Abstract (English)

Motion controllability is crucial in video synthesis. However, most previous methods are limited to single control types, and combining them often results in logical conflicts. In this paper, we propose a disentangled and unified framework, namely I2VControl, to overcome the logical conflicts. We rethink camera control, object dragging, and motion brush, reformulating all tasks into a consistent representation based on point trajectories, each managed by a dedicated formulation. Accordingly, we propose a spatial partitioning strategy, where each unit is assigned to a concomitant control category, enabling diverse control types to be dynamically orchestrated within a single synthesis pipeline without conflicts. Furthermore, we design an adapter structure that functions as a plug-in for pre-trained models and is agnostic to specific model architectures. We conduct extensive experiments, achieving excellent performance on various control tasks, and our method further facilitates user-driven creative combinations, enhancing innovation and creativity. Project page: https://wanquanf.github.io/I2VControl .

视频生成运动控制统一框架点轨迹

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。