arXiv:2509.17818cs.CV2025-09AAAI被引 28

无需训练即可精准编辑视频对象,保持时序一致性与高保真。

ContextFlow: Training-Free Video Object Editing via Adaptive Context Enrichment

  • 用高阶修正流求解器建立稳定编辑基础,提升逆向精度。
  • 通过自适应上下文增强动态融合多路径信息,避免特征替换冲突。
  • 数据驱动定位关键层,实现针对插入/替换任务的精准引导。

训练自由的视频对象编辑旨在实现精确的对象级操作,包括对象插入、替换和删除。然而,该方法在保持保真度和时序一致性方面面临重大挑战。现有方法通常针对U-Net架构设计,存在两大局限:一阶求解器导致逆向不准确,以及粗略的“硬”特征替换引发上下文冲突。这些问题在扩散变压器(DiT)中尤为突出,因先前的层选择启发式方法不适用,难以实现有效引导。为此,我们提出ContextFlow,一种基于DiT的训练自由视频对象编辑新框架。首先采用高阶修正流求解器构建稳健的编辑基础。核心是自适应上下文增强机制,通过拼接并行重建与编辑路径的键值对来丰富自注意力上下文,使模型能动态融合信息。此外,为确定增强位置,提出基于数据驱动的系统分析,结合新型引导响应度量,识别不同任务(如插入、替换)的关键层,实现针对性高效引导。大量实验表明,ContextFlow显著优于现有训练自由方法,甚至超越多个先进训练有监督方法,生成时序一致且高保真的结果。

原文摘要 · Abstract (English)

Training-free video object editing aims to achieve precise object-level manipulation, including object insertion, swapping, and deletion. However, it faces significant challenges in maintaining fidelity and temporal consistency. Existing methods, often designed for U-Net architectures, suffer from two primary limitations: inaccurate inversion due to first-order solvers, and contextual conflicts caused by crude "hard" feature replacement. These issues are more challenging in Diffusion Transformers (DiTs), where the unsuitability of prior layer-selection heuristics makes effective guidance challenging. To address these limitations, we introduce ContextFlow, a novel training-free framework for DiT-based video object editing. In detail, we first employ a high-order Rectified Flow solver to establish a robust editing foundation. The core of our framework is Adaptive Context Enrichment (for specifying what to edit), a mechanism that addresses contextual conflicts. Instead of replacing features, it enriches the self-attention context by concatenating Key-Value pairs from parallel reconstruction and editing paths, empowering the model to dynamically fuse information. Additionally, to determine where to apply this enrichment (for specifying where to edit), we propose a systematic, data-driven analysis to identify task-specific vital layers. Based on a novel Guidance Responsiveness Metric, our method pinpoints the most influential DiT blocks for different tasks (e.g., insertion, swapping), enabling targeted and highly effective guidance. Extensive experiments show that ContextFlow significantly outperforms existing training-free methods and even surpasses several state-of-the-art training-based approaches, delivering temporally coherent, high-fidelity results.

视频编辑扩散模型无训练上下文增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。