用概念移植实现精准图像编辑,无需调参即能保持视觉一致
Concept Lancet: Image Editing with Compositional Representation Transplant
- 将图像分解为视觉概念的稀疏组合,自动估算概念存在度
- 根据替换/添加/删除任务定制移植策略,提升编辑精度
- 构建15万条概念数据集,支持零样本插件式应用
扩散模型广泛用于图像编辑。现有方法通常在文本嵌入或梯度空间设计编辑方向,但面临关键挑战:过度增强会破坏视觉一致性,不足则无法完成编辑。每张源图所需编辑强度不同,手动调参成本高。为此,我们提出概念剪刀(Concept Lancet, CoLan),一种零样本、可即插即用的扩散图像编辑表示操作框架。推理时,将源图在潜在空间(文本嵌入或扩散梯度)中分解为所收集视觉概念表示的稀疏线性组合,从而精确估计各概念出现程度,指导编辑。针对替换、添加或移除任务,执行定制化概念移植过程以施加对应编辑方向。为充分建模概念空间,我们构建了包含15万条多样化视觉术语与场景描述的概念表示数据集CoLan-150K,作为潜在字典。在多个基于扩散的图像编辑基线上实验表明,集成CoLan的方法在编辑效果和一致性保持方面均达到最先进水平。
原文摘要 · Abstract (English)
Diffusion models are widely used for image editing tasks. Existing editing methods often design a representation manipulation procedure by curating an edit direction in the text embedding or score space. However, such a procedure faces a key challenge: overestimating the edit strength harms visual consistency while underestimating it fails the editing task. Notably, each source image may require a different editing strength, and it is costly to search for an appropriate strength via trial-and-error. To address this challenge, we propose Concept Lancet (CoLan), a zero-shot plug-and-play framework for principled representation manipulation in diffusion-based image editing. At inference time, we decompose the source input in the latent (text embedding or diffusion score) space as a sparse linear combination of the representations of the collected visual concepts. This allows us to accurately estimate the presence of concepts in each image, which informs the edit. Based on the editing task (replace/add/remove), we perform a customized concept transplant process to impose the corresponding editing direction. To sufficiently model the concept space, we curate a conceptual representation dataset, CoLan-150K, which contains diverse descriptions and scenarios of visual terms and phrases for the latent dictionary. Experiments on multiple diffusion-based image editing baselines show that methods equipped with CoLan achieve state-of-the-art performance in editing effectiveness and consistency preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。