arXiv:2606.28094cs.CVcs.AI2026-06

一歩で効果を考慮した物体除去、不完全なマスクにも強い。

OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

论文配图:OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
图 1 · 摘自论文原文
  • 1步の拡散モデルに境界精度を高める機構を導入
  • 4〜30倍速く、既存手法より質が高い除去結果
  • 不正確なマスクでも動作し、リアルな影・反射を再現

真实世界中的物体去除面临两大挑战:目标物体的非局部效应(如阴影和反射)难以建模,且用户提供的掩码常不准确或不完整。基于扩散模型的方法虽性能强大,但因参数量达数十亿、需数十步去噪,计算成本过高,难以用于交互式应用或边缘设备。为此,我们提出OSOR(一步物体去除),实现高效、考虑效应、对掩码鲁棒的物体去除。具体包括:(1) 基于占用引导的判别器,提供精确边界监督,支持稳定的一步扩散训练;(2) 引入α头,利用预训练扩散模型知识,以极小开销预测合理去除区域,有效应对不完整掩码;(3) 提出语义锚定验证流水线(SAVP),过滤噪声指令三元组,在大规模生成效应感知的监督信号。基于SAVP,我们构建了包含28万条验证去除对的CORNE数据集,并进一步标注AnimeEraseBench与TextEraseBench以评估复杂去除任务表现。实验表明,OSOR在感知质量上超越强大多步扩散基线,同时推理速度提升4至30倍。

原文摘要 · Abstract (English)

Real-world object removal is challenging due to two key difficulties: the target object's non-local effects, such as shadows and reflections, which are difficult to model, and the fact that user-provided masks are often inaccurate or incomplete. With billions of parameters and tens of denoising steps, diffusion-based models achieve strong removal performance at the expense of substantial computational cost, limiting their use in interactive applications and on edge devices. To address these challenges, we present OSOR (One-Step Object Removal), which simultaneously achieves efficient, effect-aware, and mask-robust object removal. Concretely, OSOR introduces: (1) an occupancy-guided discriminator for precise boundary supervision, enabling stable single-step diffusion training; (2) an alpha head that leverages knowledge from pretrained diffusion models to predict appropriate removal regions with minimal overhead, thereby handling imperfect masks; and (3) a semantic-anchored verification pipeline (SAVP) that filters noisy instruction-based triplets to produce effect-aware supervision at scale. Using SAVP, we curate CORNE, which contains 280K verified removal pairs, and further annotate AnimeEraseBench and TextEraseBench to evaluate performance on more complex removal tasks. Experiments show that OSOR surpasses strong multi-step diffusion baselines in perceptual quality while achieving $4\times$ to $30\times$ faster inference.

图像修复扩散模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。