一模型统一处理物体删除、插入和移动,可控生成光影效果。
CrimEdit: Controllable Editing for Counterfactual Object Removal, Insertion, and Movement
- 用统一模型联合训练删除与插入任务,结合无分类器指引增强效果处理。
- 单步完成物体移动,删除与插入效果更自然,无需额外训练。
- 适合需要高效、精准图像编辑的设计师与研究人员。
近期物体删除与插入工作通过处理阴影、反光等物体效应,提升了性能,采用在反事实数据集上训练的扩散模型。然而,在统一模型中应用无分类器指引来处理删除与插入任务中的物体效应,其性能影响尚未深入探索。为填补这一空白并提升复合编辑效率,我们提出CrimEdit,该方法在单个模型中联合训练删除与插入任务的嵌入表示,并利用它们在无分类器指引框架下运行——既增强了物体及其效应的删除效果,又实现了插入时物体效应的可控合成。CrimEdit还将这两个任务提示扩展至空间上不同的区域,使物体在单次去噪步骤内即可完成重新定位。通过结合两种引导技术,大量实验表明,CrimEdit在物体删除、可控效应插入及高效物体移动方面表现优异,且无需额外训练或分离删除与插入阶段。
原文摘要 · Abstract (English)
Recent works on object removal and insertion have enhanced their performance by handling object effects such as shadows and reflections, using diffusion models trained on counterfactual datasets. However, the performance impact of applying classifier-free guidance to handle object effects across removal and insertion tasks within a unified model remains largely unexplored. To address this gap and improve efficiency in composite editing, we propose CrimEdit, which jointly trains the task embeddings for removal and insertion within a single model and leverages them in a classifier-free guidance scheme -- enhancing the removal of both objects and their effects, and enabling controllable synthesis of object effects during insertion. CrimEdit also extends these two task prompts to be applied to spatially distinct regions, enabling object movement (repositioning) within a single denoising step. By employing both guidance techniques, extensive experiments show that CrimEdit achieves superior object removal, controllable effect insertion, and efficient object movement without requiring additional training or separate removal and insertion stages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。