用250万张高质量图像编辑对,让模型能精准理解复杂指令。
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
- 构建250万对高质图像编辑数据,覆盖20多种编辑类型。
- 新模型在3个基准测试中显著提升扩散模型的编辑性能。
- 适合需要精准、多样图像编辑的创意工作者使用。
基于指令的图像编辑旨在通过自然语言指令修改图像中的特定元素。然而,当前该领域模型常因训练数据质量低、编辑类型有限而难以准确执行复杂指令。我们提出AnyEdit,一个涵盖250万对高质量编辑样本的多模态指令编辑数据集,覆盖超过20种编辑类型和五个应用领域。通过初始数据多样性、自适应编辑流程与自动化结果筛选,确保数据集的丰富性与质量。基于此数据集,我们训练了新型的AnyEdit Stable Diffusion模型,采用任务感知路由与可学习任务嵌入,实现统一图像编辑。在三个基准数据集上的综合实验表明,AnyEdit持续提升基于扩散模型的编辑性能,为发展支持人类创造力的指令驱动图像编辑模型提供了可能。
原文摘要 · Abstract (English)
Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execute complex user instructions, as they are trained on low-quality data with limited editing types. We present AnyEdit, a comprehensive multi-modal instruction editing dataset, comprising 2.5 million high-quality editing pairs spanning over 20 editing types and five domains. We ensure the diversity and quality of the AnyEdit collection through three aspects: initial data diversity, adaptive editing process, and automated selection of editing results. Using the dataset, we further train a novel AnyEdit Stable Diffusion with task-aware routing and learnable task embedding for unified image editing. Comprehensive experiments on three benchmark datasets show that AnyEdit consistently boosts the performance of diffusion-based editing models. This presents prospects for developing instruction-driven image editing models that support human creativity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。