arXiv:2505.22944cs.CVcs.AI2025-05被引 37

用轨迹控制视频生成,能精准操控镜头、物体和局部运动。

ATI: Any Trajectory Instruction for Controllable Video Generation

  • 通过轻量级注入器将用户轨迹映射到潜在空间,统一控制多种运动。
  • 在多个任务上显著优于现有方法,视觉质量与可控性俱佳。
  • 适合需要精细运动控制的创作者,兼容主流视频生成模型。

我们提出一种统一的视频生成运动控制框架,通过基于轨迹的输入,无缝整合相机运动、物体级平移和细粒度局部运动。与以往分模块或任务专用设计不同,本方法通过轻量级运动注入器,将用户定义的轨迹投影到预训练图像到视频生成模型的潜在空间中。用户可指定关键点及其运动路径,以控制局部形变、整体物体运动、虚拟相机动态或其组合。注入的轨迹信号引导生成过程,实现时间一致且语义对齐的运动序列。实验表明,该框架在风格化运动效果(如运动笔刷)、动态视角变化和精确局部运动操作等多项任务中表现卓越,相比先前方法及商业方案显著提升可控性与视觉质量,同时广泛兼容各类先进视频生成主干网络。

原文摘要 · Abstract (English)

We propose a unified framework for motion control in video generation that seamlessly integrates camera movement, object-level translation, and fine-grained local motion using trajectory-based inputs. In contrast to prior methods that address these motion types through separate modules or task-specific designs, our approach offers a cohesive solution by projecting user-defined trajectories into the latent space of pre-trained image-to-video generation models via a lightweight motion injector. Users can specify keypoints and their motion paths to control localized deformations, entire object motion, virtual camera dynamics, or combinations of these. The injected trajectory signals guide the generative process to produce temporally consistent and semantically aligned motion sequences. Our framework demonstrates superior performance across multiple video motion control tasks, including stylized motion effects (e.g., motion brushes), dynamic viewpoint changes, and precise local motion manipulation. Experiments show that our method provides significantly better controllability and visual quality compared to prior approaches and commercial solutions, while remaining broadly compatible with various state-of-the-art video generation backbones. Project page: https://anytraj.github.io/.

视频生成运动控制轨迹驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。