通过对应关系提升扩散模型少步图像编辑的结构与纹理保真度
Cora: Correspondence-aware image editing using few step diffusion
- 基于语义对应关系校正噪声并设计插值注意力图
- 在姿态变化、物体添加等任务中保持结构与身份一致性
- 适合需要精细控制生成与保留平衡的视觉创作场景
图像编辑是计算机图形学、视觉和影视特效中的重要任务,近期基于扩散模型的方法实现了快速且高质量的编辑效果。然而,涉及显著结构变化(如非刚性形变、物体修改或内容生成)的编辑仍具挑战性。现有少步编辑方法常出现无关纹理或难以保持源图像关键属性(如姿态)。我们提出 Cora,通过引入对应感知的噪声修正和插值注意力图,解决上述问题。该方法利用语义对应关系对齐源图像与目标图像的纹理与结构,实现精准纹理迁移,并在必要时生成新内容。Cora 可调控生成与保留之间的平衡。大量实验表明,无论定量还是定性评估,Cora 在姿态变化、物体添加和纹理优化等多种编辑任务中均能有效保持结构、纹理与身份。用户研究进一步证实其优于现有方法。
原文摘要 · Abstract (English)
Image editing is an important task in computer graphics, vision, and VFX, with recent diffusion-based methods achieving fast and high-quality results. However, edits requiring significant structural changes, such as non-rigid deformations, object modifications, or content generation, remain challenging. Existing few step editing approaches produce artifacts such as irrelevant texture or struggle to preserve key attributes of the source image (e.g., pose). We introduce Cora, a novel editing framework that addresses these limitations by introducing correspondence-aware noise correction and interpolated attention maps. Our method aligns textures and structures between the source and target images through semantic correspondence, enabling accurate texture transfer while generating new content when necessary. Cora offers control over the balance between content generation and preservation. Extensive experiments demonstrate that, quantitatively and qualitatively, Cora excels in maintaining structure, textures, and identity across diverse edits, including pose changes, object addition, and texture refinements. User studies confirm that Cora delivers superior results, outperforming alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。