arXiv:2504.09697cs.GRcs.CV2025-04中稿 · NeurIPS被引 2

无需训练,百次编辑仍保图像质量的精准可控图像编辑流程

SPICE: A Synergistic, Precise, Iterative, and Customizable Image Editing Workflow

  • 融合扩散模型与Canny边缘控制网络,协同实现自由指令编辑
  • 百次迭代后仍保持图像质量,未编辑区域完全不变
  • 支持任意分辨率,适合艺术创作与研究复现

基于提示的模型在图像编辑任务中展现出强大的提示遵循能力,但仍难以准确执行细节丰富的编辑指令或局部修改。尤其在单次编辑后,全局图像质量常迅速下降。为此,我们提出SPICE——一种无需训练的工作流,可接受任意分辨率与长宽比,精确遵循用户需求,并在超过100次编辑步骤中持续提升图像质量,同时保持未编辑区域不变。通过协同基底扩散模型与Canny边缘ControlNet模型的优势,SPICE能稳健处理用户提供的自由形式编辑指令。在一项具有挑战性的真实图像编辑数据集上,SPICE在定量指标上超越现有最优基准,且人类标注者一致更偏好其输出。我们已开放该工作流在主流扩散模型Web UI中的实现,以支持后续研究与艺术探索。

原文摘要 · Abstract (English)

Prompt-based models have demonstrated impressive prompt-following capability at image editing tasks. However, the models still struggle with following detailed editing prompts or performing local edits. Specifically, global image quality often deteriorates immediately after a single editing step. To address these challenges, we introduce SPICE, a training-free workflow that accepts arbitrary resolutions and aspect ratios, accurately follows user requirements, and consistently improves image quality during more than 100 editing steps, while keeping the unedited regions intact. By synergizing the strengths of a base diffusion model and a Canny edge ControlNet model, SPICE robustly handles free-form editing instructions from the user. On a challenging realistic image-editing dataset, SPICE quantitatively outperforms state-of-the-art baselines and is consistently preferred by human annotators. We release the workflow implementation for popular diffusion model Web UIs to support further research and artistic exploration.

图像编辑扩散模型ControlNet无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。