无需微调即可实现精准图像编辑,保持原图结构不变。
Structure-Preserving Zero-Shot Image Editing via Stage-Wise Latent Injection in Diffusion Models
- 分阶段注入隐空间:早期注入学形,后期注入属性特征。
- 在表情迁移、纹理变换等任务上达到当前最佳效果。
- 适合需要快速适配新风格或参考图的图像编辑场景。
我们提出一种基于扩散模型的零样本图像编辑框架,统一了文本引导与参考引导方法,无需微调。通过扩散反演与时间步特定的空文本嵌入,有效保留源图像的结构完整性。采用分阶段隐空间注入策略:早期注入形状信息,后期注入属性特征,实现精细可控修改并保持全局一致性。通过参考隐空间的交叉注意力机制,实现源图与参考图之间的语义对齐。在表情迁移、纹理变换和风格融合等多种任务上的大量实验表明,该方法性能领先,展现出良好的可扩展性与适应性。
原文摘要 · Abstract (English)
We propose a diffusion-based framework for zero-shot image editing that unifies text-guided and reference-guided approaches without requiring fine-tuning. Our method leverages diffusion inversion and timestep-specific null-text embeddings to preserve the structural integrity of the source image. By introducing a stage-wise latent injection strategy-shape injection in early steps and attribute injection in later steps-we enable precise, fine-grained modifications while maintaining global consistency. Cross-attention with reference latents facilitates semantic alignment between the source and reference. Extensive experiments across expression transfer, texture transformation, and style infusion demonstrate state-of-the-art performance, confirming the method's scalability and adaptability to diverse image editing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。