提出新方法与数据集,实现高质量视频中物体及影子等效果的清除。
EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing
- 通过逆向学习机制,将物体插入作为辅助任务提升去效能力。
- 在6万对视频上训练,显著提升阴影、反射等效果的消除精度。
- 适合需要精准视频编辑与特效去除的研究者和开发者使用。
视频目标移除旨在消除动态目标及其视觉效应(如形变、阴影、反光),并恢复无缝背景。现有基于扩散模型的视频修复与目标移除方法虽能移除物体,但难以彻底清除这些效应,且背景合成不连贯。此外,缺乏系统性覆盖多种环境下的常见物体效应的数据集制约了进展。为此,我们构建了大规模视频目标移除数据集VOR,包含60K对高质量视频,每对包含一个含目标及效应的视频和对应无目标无效应版本,并提供相应物体掩码。VOR涵盖五类视觉效应,覆盖多种物体类别及复杂多目标动态场景。基于VOR,我们提出EffectErase方法,采用双向学习框架,将物体插入设为逆向辅助任务。模型引入任务感知区域引导,聚焦受影响区域,支持灵活任务切换;并设计插入-移除一致性目标,促进互补行为与效应区域及结构线索的共享定位。在VOR上训练后,EffectErase在广泛实验中表现优异,可跨多样化场景实现高质量视频效应擦除。
原文摘要 · Abstract (English)
Video object removal aims to eliminate dynamic target objects and their visual effects, such as deformation, shadows, and reflections, while restoring seamless backgrounds. Recent diffusion-based video inpainting and object removal methods can remove the objects but often struggle to erase these effects and to synthesize coherent backgrounds. Beyond method limitations, progress is further hampered by the lack of a comprehensive dataset that systematically captures common object effects across varied environments for training and evaluation. To address this, we introduce VOR (Video Object Removal), a large-scale dataset that provides diverse paired videos, each consisting of one video where the target object is present with its effects and a counterpart where the object and effects are absent, with corresponding object masks. VOR contains 60K high-quality video pairs from captured and synthetic sources, covers five effects types, and spans a wide range of object categories as well as complex, dynamic multi-object scenes. Building on VOR, we propose EffectErase, an effect-aware video object removal method that treats video object insertion as the inverse auxiliary task within a reciprocal learning scheme. The model includes task-aware region guidance that focuses learning on affected areas and enables flexible task switching. Then, an insertion-removal consistency objective that encourages complementary behaviors and shared localization of effect regions and structural cues. Trained on VOR, EffectErase achieves superior performance in extensive experiments, delivering high-quality video object effect erasing across diverse scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。