arXiv:2504.15723cs.CV2025-04被引 3

无需微调即可实现精准图像编辑,保持原图结构不变。

Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models

  • 分阶段注入隐空间:早期注入学形,后期注入属性特征。
  • 在表情迁移、纹理变换等任务上达到当前最佳效果。
  • 适合需要快速适配新风格或参考图的图像编辑场景。

我们提出一种基于扩散模型的零样本图像编辑框架,统一了文本引导与参考引导方法,无需微调。通过扩散反演与时间步特定的空文本嵌入,有效保留源图像的结构完整性。采用分阶段隐空间注入策略:早期注入形状信息,后期注入属性特征,实现精细可控修改并保持全局一致性。通过参考隐空间的交叉注意力机制,实现源图与参考图之间的语义对齐。在表情迁移、纹理变换和风格融合等多种任务上的大量实验表明,该方法性能领先,展现出良好的可扩展性与适应性。

原文摘要 · Abstract (English)

We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific null-text embeddings to preserve the structural integrity of the source image. By introducing a stage-wise latent injection strategy-shape injection in early steps and attribute injection in later steps-we enable precise, fine-grained modifications while maintaining global consistency. Cross-attention with reference latents facilitates semantic alignment between the source and reference. Extensive experiments across expression transfer, texture transformation, and style infusion demonstrate state-of-the-art performance, confirming the method's scalability and adaptability to diverse image editing scenarios.

图像编辑扩散模型零样本结构保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。