arXiv:2603.08850cs.CV2026-03

让视频生成可精准控制物体位置和运动轨迹

HECTOR: Hybrid Editable Compositional Object References for Video Generation

  • 结合静态图与动态视频双重参考,实现混合条件生成
  • 支持指定每个物体的移动路径、大小和速度变化
  • 适合需要精细控制视频元素布局与运动的创作者

现实世界视频天然呈现不同物理对象间的复杂互动,形成动态视觉组合。但现有视频生成模型多以整体方式合成场景,缺乏显式的组合控制机制。为此,我们提出HECTOR,一种支持细粒度组合控制的生成流程。与以往方法不同,HECTOR支持混合参考条件输入,可同时使用静态图像和/或动态视频进行引导。用户还能显式指定每个参考元素的运动轨迹,精确控制其位置、尺度和速度(见图1)。该设计使模型能在满足复杂时空约束的同时,保持对参考内容的高保真还原。大量实验表明,相较于现有方法,HECTOR在视觉质量、参考保留度和运动可控性方面均表现更优。

原文摘要 · Abstract (English)

Real-world videos naturally portray complex interactions among distinct physical objects, effectively forming dynamic compositions of visual elements. However, most current video generation models synthesize scenes holistically and therefore lack mechanisms for explicit compositional manipulation. To address this limitation, we propose HECTOR, a generative pipeline that enables fine-grained compositional control. In contrast to prior methods,HECTOR supports hybrid reference conditioning, allowing generation to be simultaneously guided by static images and/or dynamic videos. Moreover, users can explicitly specify the trajectory of each referenced element, precisely controlling its location, scale, and speed (see Figure1). This design allows the model to synthesize coherent videos that satisfy complex spatiotemporal constraints while preserving high-fidelity adherence to references. Extensive experiments demonstrate that HECTOR achieves superior visual quality, stronger reference preservation, and improved motion controllability compared with existing approaches.

视频生成可控生成轨迹控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。