arXiv:2503.04215cs.CV2025-03AAAI被引 3

用能量引导优化实现精准个性化图像编辑

Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models

论文配图:Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion Models
图 1 · 摘自论文原文
  • 在潜在空间中以扩散模型为能量函数引导编辑
  • 粗到精策略提升对象融合自然度与外观一致性
  • 无需训练,适合高保真个性化图像替换场景

预训练的文本驱动扩散模型极大地丰富了图像生成与编辑应用。然而,随着个性化内容编辑需求的增长,面对任意物体和复杂场景时,现有方法常将掩码误作对象形状先验,难以实现无缝融合;且普遍采用的反演噪声初始化也影响目标对象的身份一致性。为此,我们提出一种无需训练的新框架,将个性化内容编辑建模为潜在空间中编辑图像的优化问题,利用参考图文对作为条件,以扩散模型作为能量函数引导。采用粗到精策略:初期使用文本能量引导实现目标类别间的自然过渡,后期通过点对点特征级图像能量引导完成精细外观对齐。此外,引入潜在空间内容组合机制,增强整体身份一致性。大量实验表明,该方法在存在大领域差异时仍能实现优秀的目标替换效果,展现出高质量个性化图像编辑的潜力。

原文摘要 · Abstract (English)

The rapid advancement of pretrained text-driven diffusion models has significantly enriched applications in image generation and editing. However, as the demand for personalized content editing increases, new challenges emerge especially when dealing with arbitrary objects and complex scenes. Existing methods usually mistakes mask as the object shape prior, which struggle to achieve a seamless integration result. The mostly used inversion noise initialization also hinders the identity consistency towards the target object. To address these challenges, we propose a novel training-free framework that formulates personalized content editing as the optimization of edited images in the latent space, using diffusion models as the energy function guidance conditioned by reference text-image pairs. A coarse-to-fine strategy is proposed that employs text energy guidance at the early stage to achieve a natural transition toward the target class and uses point-to-point feature-level image energy guidance to perform fine-grained appearance alignment with the target object. Additionally, we introduce the latent space content composition to enhance overall identity consistency with the target. Extensive experiments demonstrate that our method excels in object replacement even with a large domain gap, highlighting its potential for high-quality, personalized image editing.

图像编辑扩散模型个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。