arXiv:2409.13037cs.CV2024-09ECCV被引 16

通过噪声稀释提升视频编辑的灵活性,支持动态运动修改。

DNI: Dilutional Noise Initialization for Diffusion Video Editing

  • 在目标区域增加额外噪声,弱化原视频结构约束。
  • 实现更贴近提示词的精准动态编辑效果。
  • 适合需要改变动作或姿态的复杂视频编辑场景。

基于文本的扩散视频编辑系统在高保真度和文本对齐方面表现优异,但仅限于风格迁移、物体叠加等刚性编辑,难以实现如运动变化这类需结构重构的非刚性编辑。原因在于扩散视频编辑中初始潜在噪声仍保留输入视频的视觉结构。本文提出噪声稀释初始化(DNI)框架,通过在待编辑区域的潜在噪声中引入更多噪声,降低原视频结构的刚性束缚,使编辑更贴近目标提示。大量实验验证了DNI的有效性。

原文摘要 · Abstract (English)

Text-based diffusion video editing systems have been successful in performing edits with high fidelity and textual alignment. However, this success is limited to rigid-type editing such as style transfer and object overlay, while preserving the original structure of the input video. This limitation stems from an initial latent noise employed in diffusion video editing systems. The diffusion video editing systems prepare initial latent noise to edit by gradually infusing Gaussian noise onto the input video. However, we observed that the visual structure of the input video still persists within this initial latent noise, thereby restricting non-rigid editing such as motion change necessitating structural modifications. To this end, this paper proposes Dilutional Noise Initialization (DNI) framework which enables editing systems to perform precise and dynamic modification including non-rigid editing. DNI introduces a concept of `noise dilution' which adds further noise to the latent noise in the region to be edited to soften the structural rigidity imposed by input video, resulting in more effective edits closer to the target prompt. Extensive experiments demonstrate the effectiveness of the DNI framework.

视频编辑扩散模型噪声控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。