用户拖拽物体即可生成多对象复杂运动视频
DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships
- 通过实体表示实现物体级控制,支持拖拽操作
- 可同时控制多个物体的复杂轨迹,保持相对空间关系
- 适合需要精细动作控制的视频创作场景
近年来,扩散模型在视频生成领域取得显著进展,可控视频生成受到广泛关注。然而,现有控制方法仍面临两大挑战:一是深度图、3D网格等控制条件对普通用户难以获取;二是难以同时驱动多个物体进行复杂运动轨迹。本文提出DragEntity,一种基于实体与位置关系的轨迹引导视频生成模型。相比以往方法,该模型具备两大优势:1)交互更友好,用户可直接拖拽图像中的实体而非像素;2)采用实体表示任意图像物体,多个物体可维持相对空间关系,从而支持不同复杂度的多轨迹同步控制。实验验证了DragEntity的有效性,展示了其在视频生成中精细控制方面的优异表现。
原文摘要 · Abstract (English)
In recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, existing control methods still face two limitations: Firstly, control conditions (such as depth maps, 3D Mesh) are difficult for ordinary users to obtain directly. Secondly, it's challenging to drive multiple objects through complex motions with multiple trajectories simultaneously. In this paper, we introduce DragEntity, a video generation model that utilizes entity representation for controlling the motion of multiple objects. Compared to previous methods, DragEntity offers two main advantages: 1) Our method is more user-friendly for interaction because it allows users to drag entities within the image rather than individual pixels. 2) We use entity representation to represent any object in the image, and multiple objects can maintain relative spatial relationships. Therefore, we allow multiple trajectories to control multiple objects in the image with different levels of complexity simultaneously. Our experiments validate the effectiveness of DragEntity, demonstrating its excellent performance in fine-grained control in video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。