arXiv:2505.22636cs.CV2025-05中稿 · CVPR被引 11

提出新方法精准移除物体及其阴影反光,避免画面失真。

Precise Object and Effect Removal with Adaptive Target-Aware Attention

  • 用自适应注意力机制分离物体与背景处理,定位更准
  • 在复杂场景中实现高保真背景还原,减少幻觉和误改
  • 构建大规模数据集OBER,支持物体与效果联合移除研究

物体移除不仅需清除目标物体,还需消除其伴随的视觉效应如阴影和反射。现有基于扩散模型的修复方法常引入伪影、产生幻觉内容、改变背景,且难以准确去除物体效应。为此,我们提出ObjectClear框架,通过自适应目标感知注意力机制,将前景移除与背景重建解耦。该设计使模型能精确识别并移除物体及其效应,同时保持背景高保真度。此外,训练所得注意力图用于推理阶段的引导融合策略,进一步提升视觉一致性。为支持训练与评估,我们构建了大规模物体-效应移除数据集OBER,包含带与不带物体-效应的配对图像及物体与效应的精确掩码,涵盖真实拍摄与模拟数据,覆盖多样物体、效应及复杂多物体场景。大量实验表明,ObjectClear优于现有方法,在挑战性场景中实现更优的物体-效应移除质量与背景保真度。

原文摘要 · Abstract (English)

Object removal requires eliminating not only the target object but also its associated visual effects such as shadows and reflections. However, diffusion-based inpainting and removal methods often introduce artifacts, hallucinate contents, alter background, and struggle to remove object effects accurately. To address these challenges, we propose ObjectClear, a novel framework that decouples foreground removal from background reconstruction via an adaptive target-aware attention mechanism. This design empowers the model to precisely localize and remove both objects and their effects while maintaining high background fidelity. Moreover, the learned attention maps are leveraged for an attention-guided fusion strategy during inference, further enhancing visual consistency. To facilitate the training and evaluation, we construct OBER, a large-scale dataset for OBject-Effect Removal, which provides paired images with and without object-effects, along with precise masks for both objects and their effects. The dataset comprises high-quality captured and simulated data, covering diverse objects, effects, and complex multi-object scenes. Extensive experiments demonstrate that ObjectClear outperforms prior methods, achieving superior object-effect removal quality and background fidelity, especially in challenging scenarios.

图像修复物体移除注意力机制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。