让图像生成视频时可灵活控制关键点运动轨迹。
FlexTraj: Image-to-Video Generation with Flexible Point Trajectory Control
- 用点的轨迹编码实现稠密与稀疏轨迹控制。
- 支持无对齐条件下的稳定生成,收敛更快。
- 适合需要精细动作控制的动画与视频编辑场景。
我们提出 FlexTraj,一种支持灵活点轨迹控制的图像到视频生成框架。FlexTraj 引入统一的基于点的运动表示,为每个点编码分割ID、时间一致的轨迹ID及可选的颜色通道以提供外观线索,从而实现稠密与稀疏轨迹控制。不同于通过令牌拼接或 ControlNet 注入轨迹条件,FlexTraj 采用高效的序列拼接方案,实现更快收敛、更强可控性与更高效推理,同时在条件未对齐情况下仍保持鲁棒性。为训练该统一点轨迹控制的视频生成器,FlexTraj 采用渐进式训练策略,逐步减少对完整监督和对齐条件的依赖。实验表明,FlexTraj 支持多粒度、非对齐条件下的轨迹控制,适用于动作克隆、拖拽式图像到视频生成、运动插值、相机重定向、灵活动作控制及网格动画等多种应用。
原文摘要 · Abstract (English)
We present FlexTraj, a framework for image-to-video generation with flexible point trajectory control. FlexTraj introduces a unified point-based motion representation that encodes each point with a segmentation ID, a temporally consistent trajectory ID, and an optional color channel for appearance cues, enabling both dense and sparse trajectory control. Instead of injecting trajectory conditions into the video generator through token concatenation or ControlNet, FlexTraj employs an efficient sequence-concatenation scheme that achieves faster convergence, stronger controllability, and more efficient inference, while maintaining robustness under unaligned conditions. To train such a unified point trajectory-controlled video generator, FlexTraj adopts an annealing training strategy that gradually reduces reliance on complete supervision and aligned condition. Experimental results demonstrate that FlexTraj enables multi-granularity, alignment-agnostic trajectory control for video generation, supporting various applications such as motion cloning, drag-based image-to-video, motion interpolation, camera redirection, flexible action control and mesh animations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。