arXiv:2603.20644cs.CV2026-03被引 2

用多智能体框架自动生成1200万张开源图像编辑数据,质量媲美商业级。

ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework

  • 构建多智能体系统,自动合成高质量图像编辑指令与结果
  • 生成1200万条数据,覆盖23类任务,在多个基准上提升超150%
  • 全开源框架,适合研究者快速扩展图像编辑数据集

基于指令的图像编辑已成为统一多模态模型的关键能力,但构建大规模、多样且高质量的编辑数据集仍面临挑战。现有数据集或依赖闭源模型标注,难以低成本扩展;或使用固定合成流程,质量与泛化能力有限。为此,我们提出ScaleEditor——一个完全开源的分层多智能体框架,用于端到端构建大规模高质量图像编辑数据集。该流程包含三部分:融入世界知识的源图像扩展、自适应多智能体编辑指令-图像合成,以及任务感知的数据质量验证机制。基于此,我们构建了目前最大的开源图像编辑数据集ScaleEdit-12M,覆盖23类任务,涵盖真实与合成多种场景。在UniWorld-V1和Bagel模型上微调后,于ImgEdit与GEdit基准性能最高提升10.4%和35.1%,在知识增强型基准RISE与KRIS-Bench上分别提升150.0%与26.5%。结果表明,开源智能体流水线可实现接近商用级别的数据质量,同时保持成本效益与可扩展性。框架与数据集将全部开源。

原文摘要 · Abstract (English)

Instruction-based image editing has emerged as a key capability for unified multimodal models (UMMs), yet constructing large-scale, diverse, and high-quality editing datasets without costly proprietary APIs remains challenging. Previous image editing datasets either rely on closed-source models for annotation, which prevents cost-effective scaling, or employ fixed synthetic editing pipelines, which suffer from limited quality and generalizability. To address these challenges, we propose ScaleEditor, a fully open-source hierarchical multi-agent framework for end-to-end construction of large-scale, high-quality image editing datasets. Our pipeline consists of three key components: source image expansion with world-knowledge infusion, adaptive multi-agent editing instruction-image synthesis, and a task-aware data quality verification mechanism. Using ScaleEditor, we curate ScaleEdit-12M, the largest open-source image editing dataset to date, spanning 23 task families across diverse real and synthetic domains. Fine-tuning UniWorld-V1 and Bagel on ScaleEdit yields consistent gains, improving performance by up to 10.4% on ImgEdit and 35.1% on GEdit for general editing benchmarks and by up to 150.0% on RISE and 26.5% on KRIS-Bench for knowledge-infused benchmarks. These results demonstrate that open-source, agentic pipelines can approach commercial-grade data quality while retaining cost-effectiveness and scalability. Both the framework and dataset will be open-sourced.

图像编辑多智能体数据集构建开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。