用视频帧数据实现物体及影子反光的无缝清除
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data

- 基于视频帧构建物体-背景对,降低数据标注成本
- 提出对象-背景引导机制,减少形状伪影和误生成内容
- 仅需物体掩码即可清除复杂场景中的物体及其视觉影响
图像修复算法在移除物体方面已取得显著进展,但仍面临两大挑战:难以处理物体产生的阴影、反光等视觉效应;容易生成形状类伪影或意外内容。本文提出 Video4Removal,一个包含超过10万张高质量样本的大规模数据集,涵盖真实物体阴影与反光。通过使用现成视觉模型从视频帧中构建物体-背景配对,大幅降低数据采集的人工成本。为避免生成形状类伪影和意外内容,提出物体-背景引导机制,利用前景物体与背景图像共同指导扩散过程,获取更丰富的上下文信息。基于上述设计,提出 OmniEraser,一种仅需物体掩码即可无缝移除物体及其视觉效应的新方法。大量实验表明,OmniEraser 在复杂真实场景中显著优于以往方法,并在动漫风格图像上展现出强泛化能力。数据集、模型与代码将公开。
原文摘要 · Abstract (English)
Inpainting algorithms have achieved remarkable progress in removing objects from images, yet still face two challenges: 1) struggle to handle the object's visual effects such as shadow and reflection; 2) easily generate shape-like artifacts and unintended content. In this paper, we propose Video4Removal, a large-scale dataset comprising over 100,000 high-quality samples with realistic object shadows and reflections. By constructing object-background pairs from video frames with off-the-shelf vision models, the labor costs of data acquisition can be significantly reduced. To avoid generating shape-like artifacts and unintended content, we propose Object-Background Guidance, an elaborated paradigm that takes both the foreground object and background images. It can guide the diffusion process to harness richer contextual information. Based on the above two designs, we present OmniEraser, a novel method that seamlessly removes objects and their visual effects using only object masks as input. Extensive experiments show that OmniEraser significantly outperforms previous methods, particularly in complex in-the-wild scenes. And it also exhibits a strong generalization ability in anime-style images. Datasets, models, and codes will be published.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。