用像素级精修解决扩散模型局部编辑的色差和接缝问题。
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
- 通过可微分的像素空间判别器放大并抑制细微色差与纹理差异。
- 训练时引入真实编辑瑕疵模拟,提升对复杂场景的鲁棒性。
- 无需修改主模型,适配多种扩散模型和编辑任务。
潜在扩散模型(LDMs)显著提升了图像修复与局部编辑的质量,但潜在空间压缩常导致像素级不一致,如色彩偏移、纹理错位及编辑边界可见接缝。现有方法如背景条件解码和像素空间调和,在实践中难以完全消除这些伪影,且在不同潜在表示或任务间泛化能力差。我们提出PixPerfect,一种像素级精修框架,可在多种LDM架构与任务中实现无缝高保真局部编辑。其核心包括:(i) 可微分的判别性像素空间,强化并抑制细微颜色与纹理差异;(ii) 全面的伪影模拟流程,使精修器在训练中暴露于真实局部编辑瑕疵;(iii) 直接的像素空间精修机制,确保对多样化潜在表示与任务的广泛适用性。在修复、物体移除与插入等基准测试中,PixPerfect显著提升感知保真度与下游编辑性能,确立了鲁棒高保真局部图像编辑的新标准。
原文摘要 · Abstract (English)
Latent Diffusion Models (LDMs) have markedly advanced the quality of image inpainting and local editing. However, the inherent latent compression often introduces pixel-level inconsistencies, such as chromatic shifts, texture mismatches, and visible seams along editing boundaries. Existing remedies, including background-conditioned latent decoding and pixel-space harmonization, usually fail to fully eliminate these artifacts in practice and do not generalize well across different latent representations or tasks. We introduce PixPerfect, a pixel-level refinement framework that delivers seamless, high-fidelity local edits across diverse LDM architectures and tasks. PixPerfect leverages (i) a differentiable discriminative pixel space that amplifies and suppresses subtle color and texture discrepancies, (ii) a comprehensive artifact simulation pipeline that exposes the refiner to realistic local editing artifacts during training, and (iii) a direct pixel-space refinement scheme that ensures broad applicability across diverse latent representations and tasks. Extensive experiments on inpainting, object removal, and insertion benchmarks demonstrate that PixPerfect substantially enhances perceptual fidelity and downstream editing performance, establishing a new standard for robust and high-fidelity localized image editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。