无需训练,提前筛选可靠生成种子,提升指令编辑成功率。
Early Timestep Zero-Shot Candidate Selection for Instruction-Guided Image Editing
- 在扩散模型早期步骤评估潜在样本背景一致性,快速筛选可靠种子。
- 实验显示成功率提升40%,计算成本降低41%至61%。
- 适合需要稳定图像编辑结果的研究者与开发者使用。
尽管扩散模型取得进展,但采样过程中的随机噪声仍导致图像生成与编辑不可靠。指令引导的图像编辑虽用户友好,却常出现背景失真等失败情况,用户需反复试错调整种子或提示词,效率低下。现有种子选择方法依赖外部验证器,适用性受限且计算开销大。本文首先建立基于多种子的图像编辑基线,利用背景一致性评分实现无监督的Best-of-N表现。在此基础上提出ELECT(Early-timestep Latent Evaluation for Candidate Selection)框架,通过在早期扩散步骤估算背景不一致程度,零样本筛选出保留背景、仅修改前景的可靠种子。ELECT以背景不一致得分排序候选种子,提前剔除不合适的样本,同时保持可编辑性。该方法可嵌入指令编辑流程,并扩展至多模态大语言模型(MLLMs),实现种子与提示词联合选择,在仅靠种子选择不足时进一步提升效果。实验表明,ELECT平均减少41%计算成本,最高达61%,显著提升背景一致性和指令遵循度,使此前失败案例成功率提升约40%,且全程无需外部监督或训练。
原文摘要 · Abstract (English)
Despite recent advances in diffusion models, achieving reliable image generation and editing remains challenging due to the inherent diversity induced by stochastic noise in the sampling process. Instruction-guided image editing with diffusion models offers user-friendly capabilities, yet editing failures, such as background distortion, frequently occur. Users often resort to trial and error, adjusting seeds or prompts to achieve satisfactory results, which is inefficient. While seed selection methods exist for Text-to-Image (T2I) generation, they depend on external verifiers, limiting applicability, and evaluating multiple seeds increases computational complexity. To address this, we first establish a multiple-seed-based image editing baseline using background consistency scores, achieving Best-of-N performance without supervision. Building on this, we introduce ELECT (Early-timestep Latent Evaluation for Candidate Selection), a zero-shot framework that selects reliable seeds by estimating background mismatches at early diffusion timesteps, identifying the seed that retains the background while modifying only the foreground. ELECT ranks seed candidates by a background inconsistency score, filtering unsuitable samples early based on background consistency while preserving editability. Beyond standalone seed selection, ELECT integrates into instruction-guided editing pipelines and extends to Multimodal Large-Language Models (MLLMs) for joint seed and prompt selection, further improving results when seed selection alone is insufficient. Experiments show that ELECT reduces computational costs (by 41 percent on average and up to 61 percent) while improving background consistency and instruction adherence, achieving around 40 percent success rates in previously failed cases - without any external supervision or training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。