用2万组数据训练出更强的指代图像编辑模型,比百万级数据训练的模型还强。
RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions
- 基于合成数据生成训练指代图像编辑模型,提升复杂场景理解能力
- 仅用2万组数据即超越百万级数据训练的基线模型
- 适合需要精准图像编辑和指代理解的研究者与开发者
尽管图像逆向与指令式编辑取得进展,现有方法在多实体复杂场景中仍表现不佳。为此,我们提出RefEdit-Bench——一个基于RefCOCO的严谨真实世界基准,即使在百万样本训练的基线模型下表现也较差。为克服此问题,我们提出RefEdit,一种基于可扩展合成数据生成管道训练的指令式编辑模型。仅使用20,000个编辑三元组训练的RefEdit,优于基于Flux/SD3且在数百万数据上训练的基线模型。跨多个基准的评估表明,该模型不仅在指代表达任务上表现优异,还在传统基准上提升性能,达到接近闭源方法的领先水平。相关数据与检查点已公开以保证可复现性。
原文摘要 · Abstract (English)
Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce RefEdit-Bench, a rigorous real-world benchmark rooted in RefCOCO, where even baselines trained on millions of samples perform poorly. To overcome this limitation, we introduce RefEdit -- an instruction-based editing model trained on our scalable synthetic data generation pipeline. Our RefEdit, trained on only 20,000 editing triplets, outperforms the Flux/SD3 model-based baselines trained on millions of data. Extensive evaluations across various benchmarks demonstrate that our model not only excels in referring expression tasks but also enhances performance on traditional benchmarks, achieving state-of-the-art results comparable to closed-source methods. We release data \& checkpoint for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。