无需训练,通过几何反演实现高保真图像编辑
InverseMeetInsert: Robust Real Image Editing via Geometric Accumulation Inversion in Guided Diffusion Models
- 利用几何累积损失增强DDIM反演,保持像素空间布局
- 结合像素级编辑与潜在空间引导,实现精准可控编辑
- 适用于局部到全局的定制化编辑,适合真实图像处理
本文提出一种名为几何逆向-交汇-像素插入(GEO)的图像编辑方法,可在无需训练的情况下实现局部到全局的定制化编辑。该方法融合文本与图像提示,核心创新包括:(i) 提出新颖的几何累积损失,增强DDIM反演以忠实保留像素空间的几何结构与布局;(ii) 设计改进的图像提示技术,将纯文本反演的像素级编辑与标准无分类器反演的潜在空间几何引导相结合。基于公开的Stable Diffusion模型,在多种图像类型和复杂提示场景下进行评估,结果表明GEO能持续生成高质量、高保真度的真实图像编辑结果。
原文摘要 · Abstract (English)
In this paper, we introduce Geometry-Inverse-Meet-Pixel-Insert, short for GEO, an exceptionally versatile image editing technique designed to cater to customized user requirements at both local and global scales. Our approach seamlessly integrates text prompts and image prompts to yield diverse and precise editing outcomes. Notably, our method operates without the need for training and is driven by two key contributions: (i) a novel geometric accumulation loss that enhances DDIM inversion to faithfully preserve pixel space geometry and layout, and (ii) an innovative boosted image prompt technique that combines pixel-level editing for text-only inversion with latent space geometry guidance for standard classifier-free reversion. Leveraging the publicly available Stable Diffusion model, our approach undergoes extensive evaluation across various image types and challenging prompt editing scenarios, consistently delivering high-fidelity editing results for real images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。