构建1000万规模高质量图像编辑数据集,突破大模型训练的尺度与质量瓶颈。
UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- 用统一验证模型替代多工具链,实现高效高质数据生成
- 推出1000万样本数据集UnicEdit-10M和新评测基准UnicBench
- 新增推理一致性与推理准确率指标,可精准诊断模型弱点
随着GPT-4o、Nano Banana、Seedream 4.0等强大多模态模型在图像编辑领域的快速发展,闭源与开源模型之间的性能差距日益扩大,主要源于大规模高质量训练数据和全面评测基准的缺乏。现有数据构建方法面临规模与质量的权衡:人工标注质量高但难扩展,自动化流水线易传播错误与噪声。为此,我们提出轻量级数据流水线,以端到端模型替代多工具链,并引入统一后验证阶段。为实现可扩展的质量控制,训练了一个70亿参数的双任务专家模型Qwen-Verify,用于高效故障检测与指令重描述。该流程生成了涵盖多种基础与复杂编辑任务的1000万样本数据集UnicEdit-10M。同时提出通用评测基准UnicBench,超越基础编辑,显式评估空间关系与知识驱动推理能力。为实现细粒度诊断,引入全新指标:非编辑一致性与推理准确率。对主流模型在UnicBench上的分析揭示其局限性,并为未来研究指明方向。
原文摘要 · Abstract (English)
With the rapid advances of powerful multimodal models such as GPT-4o, Nano Banana, and Seedream 4.0 in Image Editing, the performance gap between closed-source and open-source models is widening, primarily due to the scarcity of large-scale, high-quality training data and comprehensive benchmarks capable of diagnosing model weaknesses across diverse editing behaviors. Existing data construction methods face a scale-quality trade-off: human annotations are high-quality but not scalable, while automated pipelines suffer from error propagation and noise. To address this, we introduce a lightweight data pipeline that replaces multi-toolchains with an end-to-end model and a unified post-verification stage. For scalable quality control, we train a 7B dual-task expert model, \textbf{Qwen-Verify}, for efficient failure detection and instruction recaptioning. This pipeline yields \textbf{UnicEdit-10M}, a 10M-scale dataset spanning diverse basic and complex editing tasks. We also propose \textbf{UnicBench}, a general benchmark that extends beyond basic edits to explicitly assess spatial and knowledge-driven reasoning. To enable fine-grained diagnosis, we introduce novel metrics, including \textit{Non-edit Consistency} and \textit{Reasoning Accuracy}. Our analysis of mainstream models on UnicBench reveals their limitations and provides clear directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。