用参考图实现像素级精细编辑,无需训练即可高效操作。
PIXELS: Progressive Image Xemplar-based Editing with Latent Surgery
- 通过潜在空间手术实现多参考图渐进式编辑
- 支持局部对象修改与渐变空间变化,效果优于现有方法
- 无需微调模型,适合普通用户快速上手
当前基于语言引导的扩散模型图像编辑受限于繁琐的提示工程。本文提出PIXELS框架,利用真实世界图像样例实现无需重训练的渐进式图像编辑。该方法仅在推理阶段运行,可灵活使用任意数量参考图或跨模态提示,提供像素级或区域级细粒度控制。相比以往方法,其不依赖特定优化目标,兼容多种预训练文生图模型,且避免高昂计算开销。实验表明,该方法在定量指标和人工评估中均表现优异,支持选择性对象修改与渐进式空间变化,显著提升编辑自由度与质量。通过降低专业编辑门槛,使开源生成模型也能实现高质量图像修改。
原文摘要 · Abstract (English)
Recent advancements in language-guided diffusion models for image editing are often bottle-necked by cumbersome prompt engineering to precisely articulate desired changes. An intuitive alternative calls on guidance from in-the-wild image exemplars to help users bring their imagined edits to life. Contemporary exemplar-based editing methods shy away from leveraging the rich latent space learnt by pre-existing large text-to-image (TTI) models and fall back on training with curated objective functions to achieve the task. Though somewhat effective, this demands significant computational resources and lacks compatibility with diverse base models and arbitrary exemplar count. On further investigation, we also find that these techniques restrict user control to only applying uniform global changes over the entire edited region. In this paper, we introduce a novel framework for progressive exemplar-driven editing with off-the-shelf diffusion models, dubbed PIXELS, to enable customization by providing granular control over edits, allowing adjustments at the pixel or region level. Our method operates solely during inference to facilitate imitative editing, enabling users to draw inspiration from a dynamic number of reference images, or multimodal prompts, and progressively incorporate all the desired changes without retraining or fine-tuning existing TTI models. This capability of fine-grained control opens up a range of new possibilities, including selective modification of individual objects and specifying gradual spatial changes. We demonstrate that PIXELS delivers high-quality edits efficiently, leading to a notable improvement in quantitative metrics as well as human evaluation. By making high-quality image editing more accessible, PIXELS has the potential to enable professional-grade edits to a wider audience with the ease of using any open-source image generation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。