用框选精准指导图像编辑,保持背景不变
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance
- 通过多层级框选引导扩散模型定位目标区域
- 在1000张图的评测中显著优于现有开源模型
- 适合需要精确控制编辑区域的研究与应用
基于扩散模型的图像编辑已取得显著进展,但传统方法依赖自然语言提示,难以精确定位目标对象,导致全局重生成时背景一致性差。为此,我们提出利用边界框作为视觉引导,明确指定编辑区域,从而提升定位精度并保持背景稳定。为此,我们设计了细粒度框选注入方法FineEdit,使模型更有效利用空间信息。同时构建了包含120万对图像的FineEdit-1.2M数据集,以及涵盖10类主题共1000张图像的FineEdit-Bench评测基准。实验表明,该模型在指令遵循和背景保持方面显著优于Qwen-Image-Edit、LongCat-Image-Edit等先进开源模型;在GEdit和ImgEdit Bench等公开基准上也展现出更强泛化能力与鲁棒性。
原文摘要 · Abstract (English)
Diffusion-based image editing models have achieved significant progress in real world applications. However, conventional models typically rely on natural language prompts, which often lack the precision required to localize target objects. Consequently, these models struggle to maintain background consistency due to their global image regeneration paradigm. Recognizing that visual cues provide an intuitive means for users to highlight specific areas of interest, we utilize bounding boxes as guidance to explicitly define the editing target. This approach ensures that the diffusion model can accurately localize the target while preserving background consistency. To achieve this, we propose FineEdit, a multi-level bounding box injection method that enables the model to utilize spatial conditions more effectively. To support this high precision guidance, we present FineEdit-1.2M, a large scale, fine-grained dataset comprising 1.2 million image editing pairs with precise bounding box annotations. Furthermore, we construct a comprehensive benchmark, termed FineEdit-Bench, which includes 1,000 images across 10 subjects to effectively evaluate region based editing capabilities. Evaluations on FineEdit-Bench demonstrate that our model significantly outperforms state-of-the-art open-source models (e.g., Qwen-Image-Edit and LongCat-Image-Edit) in instruction compliance and background preservation. Further assessments on open benchmarks (GEdit and ImgEdit Bench) confirm its superior generalization and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。