提出GIDE框架,实现无需训练的高精度图像编辑
GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
- 通过离散噪声反演技术,精准捕捉令牌空间中的噪声模式
- 在GIDE-Bench上提升语义正确性51.83%、感知质量50.39%
- 支持文本/点/框等多种指令,保留未编辑背景
尽管扩散大语言模型(DLLMs)在多模态生成方面表现卓越,但实现精确且无需训练的图像编辑仍是开放挑战。与连续扩散模型不同,DLLMs固有的离散标记化阻碍了标准噪声反演技术的应用,常导致编辑过程中结构退化。本文提出GIDE(基于定位的DLLM图像编辑噪声反演),一个统一框架,引入新颖的离散噪声反演机制,准确捕捉离散令牌空间中的潜在噪声模式,确保高保真重建。将编辑流程分解为定位、反演和优化三个阶段,使GIDE支持多种编辑指令(文本、点、框)和操作,同时严格保留未编辑背景。此外,为克服现有单步评估协议局限,构建了包含805种组合式编辑场景的严谨基准GIDE-Bench,涵盖多样化多模态输入。在GIDE-Bench上的大量实验表明,GIDE显著优于先前无需训练的方法,语义正确性提升51.83%,感知质量提升50.39%。在ImgEdit-Bench上的额外评估也证实其广泛适用性,优于训练过的基线模型,并在真实感一致性方面达到领先模型水平。
原文摘要 · Abstract (English)
While Diffusion Large Language Models (DLLMs) have demonstrated remarkable capabilities in multi-modal generation, performing precise, training-free image editing remains an open challenge. Unlike continuous diffusion models, the discrete tokenization inherent in DLLMs hinders the application of standard noise inversion techniques, often leading to structural degradation during editing. In this paper, we introduce GIDE (Grounded Inversion for DLLM Image Editing), a unified framework designed to bridge this gap. GIDE incorporates a novel Discrete Noise Inversion mechanism that accurately captures latent noise patterns within the discrete token space, ensuring high-fidelity reconstruction. We then decompose the editing pipeline into grounding, inversion, and refinement stages. This design enables GIDE supporting various editing instructions (text, point and box) and operations while strictly preserving the unedited background. Furthermore, to overcome the limitations of existing single-step evaluation protocols, we introduce GIDE-Bench, a rigorous benchmark comprising 805 compositional editing scenarios guided by diverse multi-modal inputs. Extensive experiments on GIDE-Bench demonstrate that GIDE significantly outperforms prior training-free methods, improving Semantic Correctness by 51.83% and Perceptual Quality by 50.39%. Additional evaluations on ImgEdit-Bench confirm its broad applicability, demonstrating consistent gains over trained baselines and yielding photorealistic consistency on par with leading models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。