arXiv:2607.14976cs.CV2026-07中稿 · ECCV被引 1

一歩で動画オブジェクト除去を実現、リアルさと速度の両立

From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting

论文配图:From Draft to Draft-Free: One-Step Video Object Removal via Privileged Distillation and Fast Planting
图 1 · 摘自论文原文
  • 教師モデルが粗い除去結果を複数ステップで改善し、学生モデルに能力を蒸留
  • 1秒で1動画のノイズ除去可能、複数メトリクスで最新性能
  • 自己ガイドの疑似草案生成で草案不要な一歩モデルを実現

视频对象移除是视频编辑中的基础但极具挑战性的任务。现有方法主要分为两类:基于光流或注意力机制的传统方法常引入明显伪影,结果不自然;而扩散模型虽提升视觉真实感,但需多步去噪,实用性受限。为此,我们提出从草稿到无草稿(D2DF)框架,将多步精炼粗略移除结果的能力蒸馏为单步生成模型。教师模型通过多步将低质量移除结果(“草稿”)优化为高保真视频;通过先验特权一致性蒸馏(PPCD),将此能力转移到仅依赖草稿的一步学生模型。为消除对草稿的依赖,我们引入基于时序掩码变压器的自引导快速播种(SGFP)模块,在潜在空间中自主生成场景一致的伪草稿,实现完全无草稿的一步模型。大量实验表明,无论有无草稿版本均在多个指标上达到最先进水平,优于传统与多步生成方法,在质量和效率上均有显著提升。单个视频的去噪过程仅需约1秒。

原文摘要 · Abstract (English)

Video object removal is a fundamental yet challenging task in video editing. Despite recent progress, existing methods typically fall into two categories. Traditional approaches based on optical flow or attention mechanisms often introduce noticeable artifacts and yield unnatural results. In contrast, diffusion-based methods improve visual realism but demand multiple denoising steps, limiting their practicality. To address these issues, we propose From-Draft-to-Draft-Free (D2DF), a framework that distills the ability of transforming coarse drafts into refined videos into a one-step video generation model. Within D2DF, a teacher model is trained to refine low-quality removal results ("drafts") into high-fidelity videos by multiple steps. Then, through Prior-Privileged Consistency Distillation (PPCD), we distill this capability into a student model that performs one-step removal conditioned on the draft. To eliminate draft dependency, we introduce a Self-Guided Fast Planting (SGFP) module based on our Temporal Masked Transformer that autonomously generates scene-consistent pseudo-drafts in latent space, enabling a fully draft-free one-step model. Extensive experiments show that both draft-conditioned and draft-free versions achieve state-of-the-art performance on multiple metrics, surpassing traditional and multi-step generative methods in both quality and efficiency. The denoising process for a single video takes only about 1 second.

视频生成扩散模型去噪一歩モデル

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。