通过轨迹注意力实现视频相机运动的精准控制
Trajectory Attention for Fine-grained Video Motion Control
- 在传统时间注意力外增加轨迹注意力分支,融合像素运动路径信息
- 在部分轨迹可用时仍保持高精度与长程一致性,生成质量不降
- 适用于相机运动控制与首帧引导视频编辑,适合需要空间时间一致性的场景
视频生成的最新进展主要由视频扩散模型推动,其中相机运动控制已成为生成视角定制化视觉内容的关键挑战。本文提出轨迹注意力(trajectory attention),一种沿已有像素轨迹执行注意力的新方法,实现细粒度相机运动控制。与现有方法常导致输出不精确或忽略时序相关性不同,该方法具备更强归纳偏置,可将轨迹信息无缝注入生成过程。重要的是,轨迹注意力被建模为独立于传统时间注意力的辅助分支,使两者协同工作,在轨迹仅部分可用时仍能保证精确运动控制与新内容生成能力。在图像与视频相机运动控制任务上的实验表明,该方法显著提升精度与长程一致性,同时保持高质量生成。此外,本方法可扩展至其他视频运动控制任务,如首帧引导视频编辑,在大时空范围内表现出优异的内容一致性。
原文摘要 · Abstract (English)
Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a novel approach that performs attention along available pixel trajectories for fine-grained camera motion control. Unlike existing methods that often yield imprecise outputs or neglect temporal correlations, our approach possesses a stronger inductive bias that seamlessly injects trajectory information into the video generation process. Importantly, our approach models trajectory attention as an auxiliary branch alongside traditional temporal attention. This design enables the original temporal attention and the trajectory attention to work in synergy, ensuring both precise motion control and new content generation capability, which is critical when the trajectory is only partially available. Experiments on camera motion control for images and videos demonstrate significant improvements in precision and long-range consistency while maintaining high-quality generation. Furthermore, we show that our approach can be extended to other video motion control tasks, such as first-frame-guided video editing, where it excels in maintaining content consistency over large spatial and temporal ranges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。