arXiv:2603.27555cs.CV2026-03

无需微调或提示,一键消除图像中多个物体且保持背景自然。

PANDORA: Pixel-wise Attention Dissolution and Latent Guidance for Zero-Shot Object Removal

  • 通过像素级注意力溶解,直接清除目标物体的特征响应。
  • 在单次推理中实现多物体精准、非刚性移除,视觉质量优于现有方法。
  • 适合需要快速、无依赖图像修复的设计师与开发者使用。

从自然图像中移除物体极具挑战,主要难点在于合成语义一致的内容同时保持背景完整性。现有方法常依赖微调、提示工程或推理时优化,但仍存在纹理不一致、刚性伪影、前景-背景解耦弱以及多物体移除扩展性差等问题。本文提出一种全新的零样本物体移除框架 PANDORA,直接作用于预训练的文本到图像扩散模型,无需微调、提示或优化。提出像素级注意力溶解机制,通过抑制掩码像素最相关的注意力键,有效切断物体在自注意力流中的影响,使背景上下文主导重建过程。进一步引入局部注意力解耦引导,引导去噪过程向有利于干净移除的潜在空间演化。两者结合实现单次推理下的精确、非刚性、免提示、可扩展的多物体擦除。实验表明,该方法在视觉保真度和语义合理性上均优于当前最优方法。

原文摘要 · Abstract (English)

Removing objects from natural images is challenging due to difficulty of synthesizing semantically coherent content while preserving background integrity. Existing methods often rely on fine-tuning, prompt engineering, or inference-time optimization, yet still suffer from texture inconsistency, rigid artifacts, weak foreground-background disentanglement, and poor scalability for multi-object removal. We propose a novel zero-shot object removal framework, namely PANDORA, that operates directly on pre-trained text-to-image diffusion models, requiring no fine-tuning, prompts, or optimization. We propose Pixel-wise Attention Dissolution to remove object by nullifying the most correlated attention keys for masked pixels, effectively eliminating the object from self-attention flow and allowing background context to dominate reconstruction. We further introduce Localized Attentional Disentanglement Guidance to steer denoising toward latent manifolds favorable to clean object removal. Together, these components enable precise, non-rigid, prompt-free, and scalable multi-object erasure in a single pass. Experiments demonstrate superior visual fidelity and semantic plausibility compared to state-of-the-art methods. The project page is available at https://vdkhoi20.github.io/PANDORA.

图像修复扩散模型零样本物体移除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。