用草图控制视频生成与编辑,实现精准布局与运动调控。
SketchVideo: Sketch-based Video Generation and Editing
- 通过草图块预测跳过模块的残差特征,实现高效控制。
- 仅需1-2个关键帧草图即可全局控制视频空间与动态。
- 支持真实或合成视频的细粒度编辑,保持时空一致性。
基于文本或图像的视频生成与编辑已取得显著进展,但仅依赖文本难以精确控制整体布局与几何细节,且图像难以支持运动控制与局部修改。本文提出SketchVideo,基于DiT视频生成模型,设计一种内存高效的草图控制结构,通过草图块预测跳过DiT模块的残差特征。用户可在任意时间点的关键帧上绘制草图,实现便捷交互。为将稀疏的草图条件传播至所有帧,提出跨帧注意力机制,分析关键帧与各视频帧间关系。针对草图化视频编辑,设计额外的视频插入模块,确保新内容与原视频在空间特征和动态运动上的一致性。推理时采用潜在融合策略,精确保留未编辑区域。大量实验表明,SketchVideo在可控视频生成与编辑任务中表现优异。
原文摘要 · Abstract (English)
Video generation and editing conditioned on text prompts or images have undergone significant advancements. However, challenges remain in accurately controlling global layout and geometry details solely by texts, and supporting motion control and local modification through images. In this paper, we aim to achieve sketch-based spatial and motion control for video generation and support fine-grained editing of real or synthetic videos. Based on the DiT video generation model, we propose a memory-efficient control structure with sketch control blocks that predict residual features of skipped DiT blocks. Sketches are drawn on one or two keyframes (at arbitrary time points) for easy interaction. To propagate such temporally sparse sketch conditions across all frames, we propose an inter-frame attention mechanism to analyze the relationship between the keyframes and each video frame. For sketch-based video editing, we design an additional video insertion module that maintains consistency between the newly edited content and the original video's spatial feature and dynamic motion. During inference, we use latent fusion for the accurate preservation of unedited regions. Extensive experiments demonstrate that our SketchVideo achieves superior performance in controllable video generation and editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。