构建10万+高质量图像编辑数据集,提升复杂指令编辑能力
MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
- 用双模态大模型生成适配视觉的编辑指令与高保真图像
- 涵盖18类非风格迁移、38类风格迁移,覆盖复杂语义操作
- 适合研究复杂图像编辑、多任务泛化能力的学者使用
当前指令式图像编辑(IBIE)方法在复杂任务上表现受限,因现有数据集的编辑类型和样本数量有限,且传统构建方式常含噪声图文对,引入偏差并制约模型在复杂场景下的能力。为此,我们提出MultiEdit,一个包含超过10.7万条高质量图像编辑样本的综合性数据集。该数据集涵盖6类挑战性编辑任务,整合18种非风格迁移编辑类型和38种风格迁移操作,覆盖从复杂风格迁移至人像参考编辑、图像内文本编辑等高级语义操作。我们采用创新的数据集构建流程,利用两个多模态大语言模型(MLLMs)分别生成视觉自适应的编辑指令和生成高保真编辑图像。大量实验表明,使用MultiEdit-Train集微调开源基础模型,显著提升其在新提出的MultiEdit-Test基准上的复杂编辑任务表现,同时有效保留其在标准基准上的能力。我们认为MultiEdit为推进更丰富、更具挑战性的IBIE研究提供了宝贵资源。数据集已公开于https://huggingface.co/datasets/inclusionAI/MultiEdit。
原文摘要 · Abstract (English)
Current instruction-based image editing (IBIE) methods struggle with challenging editing tasks, as both editing types and sample counts of existing datasets are limited. Moreover, traditional dataset construction often contains noisy image-caption pairs, which may introduce biases and limit model capabilities in complex editing scenarios. To address these limitations, we introduce MultiEdit, a comprehensive dataset featuring over 107K high-quality image editing samples. It encompasses 6 challenging editing tasks through a diverse collection of 18 non-style-transfer editing types and 38 style transfer operations, covering a spectrum from sophisticated style transfer to complex semantic operations like person reference editing and in-image text editing. We employ a novel dataset construction pipeline that utilizes two multi-modal large language models (MLLMs) to generate visual-adaptive editing instructions and produce high-fidelity edited images, respectively. Extensive experiments demonstrate that fine-tuning foundational open-source models with our MultiEdit-Train set substantially improves models' performance on sophisticated editing tasks in our proposed MultiEdit-Test benchmark, while effectively preserving their capabilities on the standard editing benchmark. We believe MultiEdit provides a valuable resource for advancing research into more diverse and challenging IBIE capabilities. Our dataset is available at https://huggingface.co/datasets/inclusionAI/MultiEdit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。