arXiv:2411.11045cs.CV2024-11被引 31

解决视频编辑中动作与内容不一致的问题,提升生成稳定性。

StableV2V: Stablizing Shape Consistency in Video-to-Video Editing

  • 分步编辑:先改首帧,再对齐动作与用户提示,最后传播到所有帧。
  • 在DAVIS-Edit基准上优于现有方法,视觉一致性更强。
  • 适合需要精准控制视频内容变化的创作者和研究人员。

生成式AI的进展极大推动了内容创作与编辑,当前研究将这一趋势延伸至视频编辑。然而,现有方法主要从源视频迁移运动模式,常因动作与编辑内容缺乏对齐,导致结果与用户提示不一致。为此,本文提出一种形状一致的视频编辑方法StableV2V。该方法将编辑流程分解为多个步骤:首先编辑首帧,接着建立传递动作与用户提示之间的对齐关系,最后基于此对齐将编辑内容传播至其余所有帧。此外,我们构建了一个名为DAVIS-Edit的测试基准,用于全面评估视频编辑性能,涵盖多种提示类型与难度。实验结果与分析表明,本方法在性能、视觉一致性及推理效率方面均优于现有最先进方法。

原文摘要 · Abstract (English)

Recent advancements of generative AI have significantly promoted content creation and editing, where prevailing studies further extend this exciting progress to video editing. In doing so, these studies mainly transfer the inherent motion patterns from the source videos to the edited ones, where results with inferior consistency to user prompts are often observed, due to the lack of particular alignments between the delivered motions and edited contents. To address this limitation, we present a shape-consistent video editing method, namely StableV2V, in this paper. Our method decomposes the entire editing pipeline into several sequential procedures, where it edits the first video frame, then establishes an alignment between the delivered motions and user prompts, and eventually propagates the edited contents to all other frames based on such alignment. Furthermore, we curate a testing benchmark, namely DAVIS-Edit, for a comprehensive evaluation of video editing, considering various types of prompts and difficulties. Experimental results and analyses illustrate the outperforming performance, visual consistency, and inference efficiency of our method compared to existing state-of-the-art studies.

视频编辑一致性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。