让视频物体删除更彻底,连影子反光都一并清除
From Understanding to Erasing: Towards Complete and Stable Video Object Removal

- 通过理解物体与影子等效应的关系,精准定位影响区域
- 在多个数据集上实现无残留、时序一致的删除效果
- 适合需要高质量视频编辑的影视制作与自动驾驶场景
视频物体删除旨在擦除目标物体并重建视觉合理且时序连贯的内容。然而,目标物体常引发阴影、反光、光照变化等超出标注掩码范围的影响,导致传统掩码引导的修复方法产生可见残余。为此,我们提出一种感知侧效的物体删除方法,将其建模为理解引导的过程,整合物体-效应关系、受影响区域定位与上下文感知重建。具体而言,引入物体诱导关系蒸馏,将预训练视觉基础模型中的物体-效应级别关系迁移至视频扩散模型;设计物体感知帧级上下文交叉注意力,融合目标物体语义与每帧背景上下文以实现删除与重建;提出注意力引导区域定位,生成目标物体及其影响区域的软空间先验。在多个基准上的大量实验表明,本方法在物体及效应移除完整性、侧效抑制和时序一致性方面均优于现有方法。
原文摘要 · Abstract (English)
Video object removal aims to erase target objects while reconstructing visually plausible and temporally coherent content. However, target objects often induce shadows, reflections, illumination changes and other effects that extend beyond the provided mask, making conventional mask-conditioned completion prone to visible residuals. We therefore formulate side-effect-aware object removal as an understanding-guided process that integrates object--effect relations, affected-region localization, and context-aware reconstruction. Specifically, we introduce Object-Induced Relation Distillation to transfer token-level object--effect relations from a pretrained vision foundation model to the video diffusion model. We then design Object-aware Framewise Context Cross-Attention to combine target-object semantics with per-frame background context for removal and reconstruction, and propose Attention-guided Region Localization to derive a soft spatial prior over the target object and its affected regions. Extensive experiments across multiple benchmarks demonstrate that our method achieves more complete object-and-effect removal and outperforms existing approaches in removal quality, side-effect suppression, and temporal consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。