arXiv:2605.06535cs.CVcs.AI2026-05

解决视频背景替换中真实感不足的问题,提出新数据集与评估基准。

Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance

论文配图:Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance
图 1 · 摘自论文原文
  • 分离前景与背景引导生成,提升合成质量
  • 构建14万对视频数据,覆盖五类常见场景变换
  • 适合视频编辑、影视制作等创意应用研究者

近年来,开源项目如Senorita-2M推动了基于自然语言指令的视频编辑发展。然而现有公开数据集多聚焦局部编辑或风格迁移,主要保留原场景结构,易于扩展。相比之下,背景替换作为影视制作和广告创作中的核心任务,需生成全新且时序一致的背景,同时保持前景与背景的准确交互,大规模数据生成难度高。因此,该任务因高质量训练数据稀缺而长期被忽视,导致当前最优模型(如Kiwi-Edit)表现不佳——主要依赖的开源数据集OpenVE-3M常生成静态、不自然的背景。本文发现此问题源于数据合成阶段背景引导不够精确。为此,我们设计了一套可扩展的解耦生成管道,通过严格质量过滤实现前景与背景引导的独立生成。基于该管道,我们推出Sparkle数据集(约14万对视频,涵盖五类常见背景变换主题),以及Sparkle-Bench——目前最大的背景替换专用评估基准。实验表明,使用本数据集训练的模型在OpenVE-Bench与Sparkle-Bench上均显著优于所有现有基线。相关数据集、基准与模型已开源:https://showlab.github.io/Sparkle/

原文摘要 · Abstract (English)

In recent years, open-source efforts like Senorita-2M have propelled video editing toward natural language instruction. However, current publicly available datasets predominantly focus on local editing or style transfer, which largely preserve the original scene structure and are easier to scale. In contrast, Background Replacement, a task central to creative applications such as film production and advertising, requires synthesizing entirely new, temporally consistent scenes while maintaining accurate foreground-background interactions, making large-scale data generation significantly more challenging. Consequently, this complex task remains largely underexplored due to a scarcity of high-quality training data. This gap is evident in poorly performing state-of-the-art models, e.g., Kiwi-Edit, because the primary open-source dataset that contains this task, i.e., OpenVE-3M, frequently produces static, unnatural backgrounds. In this paper, we trace this quality degradation to a lack of precise background guidance during data synthesis. Accordingly, we design a scalable pipeline that generates foreground and background guidance in a decoupled manner with strict quality filtering. Building on this pipeline, we introduce Sparkle, a dataset of ~140K video pairs spanning five common background-change themes, alongside Sparkle-Bench, the largest evaluation benchmark tailored for background replacement to date. Experiments demonstrate that our dataset and the model trained on it achieve substantially better performance than all existing baselines on both OpenVE-Bench and Sparkle-Bench. Our proposed dataset, benchmark, and model are fully open-sourced at https://showlab.github.io/Sparkle/.

视频生成背景替换数据集自然语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。