arXiv:2410.12526cs.CV2024-10IJCAI被引 4

通过抽象概念对提升视频编辑稳定性,实现更自然的文本驱动视频修改。

Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing

  • 用抽象概念对替代关键词,避免干扰注意力机制。
  • 双先验监督机制显著提升视频稳定性和保真度。
  • 适合需要精细控制风格和动态一致性的视频编辑场景。

基于生成式扩散模型的文本驱动视频编辑因其广泛应用前景受到关注。然而,现有方法受限于预训练中提供的有限词嵌入,难以针对开放概念及特定属性进行细致编辑。直接修改目标提示中的关键词常导致注意力机制意外扰动。为此,本文提出一种改进的概念增强型视频编辑方法,通过设计抽象概念对,灵活生成多样且稳定的视频。该框架包含概念增强型文本反转与双先验监督机制。前者实现即插即用的稳定扩散引导,有效捕捉目标属性以获得更具风格化的结果;后者显著提升视频稳定性与保真度。全面评估表明,该方法生成的视频更稳定、更逼真,优于当前最先进方法。

原文摘要 · Abstract (English)

Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training, which hinders nuanced editing targeting open concepts with specific attributes. Directly altering the keywords in target prompts often results in unintended disruptions to the attention mechanisms. To achieve more flexible editing easily, this work proposes an improved concept-augmented video editing approach that generates diverse and stable target videos flexibly by devising abstract conceptual pairs. Specifically, the framework involves concept-augmented textual inversion and a dual prior supervision mechanism. The former enables plug-and-play guidance of stable diffusion for video editing, effectively capturing target attributes for more stylized results. The dual prior supervision mechanism significantly enhances video stability and fidelity. Comprehensive evaluations demonstrate that our approach generates more stable and lifelike videos, outperforming state-of-the-art methods.

视频编辑扩散模型概念增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。