arXiv:2505.23134cs.CVcs.AI2025-05被引 1

用参考帧实现零样本视频外观编辑,提升控制精度与一致性

Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing

  • 先编辑参考帧再传播外观,分离编辑过程
  • 在大运动场景下比光流更稳定,提升时序一致性
  • 适合需要精细控制视频外观的创作者

根据用户需求进行外观编辑是视频编辑的核心任务。现有文本引导方法常因意图模糊而难以实现细粒度控制。本文提出{Zero-to-Hero},采用参考帧驱动的视频编辑策略,将编辑过程拆解为两步:先生成满足用户需求的锚帧作为参考,再一致地将该外观传播至其他帧。利用原始帧间对应关系引导注意力机制,在内存友好的生成模型中表现优于传统光流或时序模块,尤其适用于大运动物体。该方法具备强零样本初始化能力,确保准确性和时序一致性。但干预注意力机制导致图像退化,出现过饱和色彩与未知模糊。为此,我们从零阶段(Zero-Stage)进入英雄阶段(Hero-Stage),构建条件生成模型以实现视频修复。为精准评估外观一致性,使用Blender构建含多外观的视频数据集,支持细粒度、确定性评测。实验表明,本方法在最佳基线基础上提升2.6 dB PSNR。项目主页见https://github.com/Tonniia/Zero2Hero。

原文摘要 · Abstract (English)

Appearance editing according to user needs is a pivotal task in video editing. Existing text-guided methods often lead to ambiguities regarding user intentions and restrict fine-grained control over editing specific aspects of objects. To overcome these limitations, this paper introduces a novel approach named {Zero-to-Hero}, which focuses on reference-based video editing that disentangles the editing process into two distinct problems. It achieves this by first editing an anchor frame to satisfy user requirements as a reference image and then consistently propagating its appearance across other frames. We leverage correspondence within the original frames to guide the attention mechanism, which is more robust than previously proposed optical flow or temporal modules in memory-friendly video generative models, especially when dealing with objects exhibiting large motions. It offers a solid ZERO-shot initialization that ensures both accuracy and temporal consistency. However, intervention in the attention mechanism results in compounded imaging degradation with over-saturated colors and unknown blurring issues. Starting from Zero-Stage, our Hero-Stage Holistically learns a conditional generative model for vidEo RestOration. To accurately evaluate the consistency of the appearance, we construct a set of videos with multiple appearances using Blender, enabling a fine-grained and deterministic evaluation. Our method outperforms the best-performing baseline with a PSNR improvement of 2.6 dB. The project page is at https://github.com/Tonniia/Zero2Hero.

视频编辑参考帧零样本外观修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。