用文字标注的轨迹精准控制视频中物体的运动与外观。
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
- 以文字描述的轨迹作为运动控制信号,实现局部精细调控。
- 在两百万高质量视频上训练,显著提升文本对齐精度与运动可控性。
- 适合需要精确控制多物体运动轨迹的视频生成任务。
文本到视频生成在视觉保真度方面进展迅速,但标准方法在控制场景中主体构图方面仍能力有限。先前工作表明,添加局部文本控制信号(如边界框或分割掩码)有所帮助,但在复杂场景和多对象设置中表现不佳,精度有限且随着可控制对象数量增加,轨迹与视觉实体之间的对应关系变得模糊。本文提出文本接地轨迹(TGT)框架,通过与局部文本描述配对的轨迹来引导视频生成。我们设计了位置感知交叉注意力(LACA)以整合这些信号,并采用双分类器引导(dual-CFG)方案分别调节局部和全局文本指导。此外,我们开发了一个数据处理管道,生成带有被追踪实体局部描述的轨迹,并标注了两百万个高质量视频片段用于训练TGT。这些组件共同使TGT能够使用点轨迹作为直观的运动控制手柄,将每个轨迹与文本配对,以同时控制外观和运动。大量实验表明,相较于以往方法,TGT在视觉质量、文本对齐准确性和运动可控性方面均有显著提升。
原文摘要 · Abstract (English)
Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However, these methods struggle in complex scenarios and degrade in multi-object settings, offering limited precision and lacking a clear correspondence between individual trajectories and visual entities as the number of controllable objects increases. We introduce Text-Grounded Trajectories (TGT), a framework that conditions video generation on trajectories paired with localized text descriptions. We propose Location-Aware Cross-Attention (LACA) to integrate these signals and adopt a dual-CFG scheme to separately modulate local and global text guidance. In addition, we develop a data processing pipeline that produces trajectories with localized descriptions of tracked entities, and we annotate two million high quality video clips to train TGT. Together, these components enable TGT to use point trajectories as intuitive motion handles, pairing each trajectory with text to control both appearance and motion. Extensive experiments show that TGT achieves higher visual quality, more accurate text alignment, and improved motion controllability compared with prior approaches. Website: https://textgroundedtraj.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。