自动化挖掘高质量图像编辑三元组,无需人工标注。
NoHumansRequired: Autonomous High-Quality Image Editing Triplet Mining
- 用自研验证器自动评估指令遵循与美学,无需分割模型。
- 生成720万条高保真三元组,数据量提升2.6倍。
- 适合图像编辑、生成模型训练研究者使用。
生成式模型的进展使图像编辑助手能无须用户干预即可执行自然语言指令。其监督训练需数百万个三元组(原图、指令、编辑图),但精确挖掘困难:每处修改必须仅作用于指定区域,保持风格一致、物理合理且视觉美观。现有自动化质量评估方法不足,限制了大规模应用。本文提出一套全自动模块化流程,可在多领域、多分辨率、复杂指令和多样风格下挖掘高质量三元组。系统基于公开生成模型,全程无需人工干预,采用任务调优的Gemini验证器直接评分指令遵循性与美学,无需依赖分割或定位模型。通过反演与组合式自举方法,数据集规模扩大约2.6倍,支持大规模高保真训练。该方法消除重复标注负担,实现无人工标注的大规模训练。为推动研究普及,我们发布NHR-Edit数据集,包含720万条工业级筛选的高质量三元组,基于数百万次引导生成与验证筛选而成,并分析各阶段留存率,提供不同模型堆栈的计算成本估算框架。在最大跨数据集评估中,性能超越所有公开替代方案。同时发布Bagel-NHR-Edit,一个微调后的Bagel模型,达到当前最佳指标。
原文摘要 · Abstract (English)
Recent advances in generative modeling enable image editing assistants that follow natural language instructions without additional user input. Their supervised training requires millions of triplets (original image, instruction, edited image), yet mining pixel-accurate examples is hard. Each edit must affect only prompt-specified regions, preserve stylistic coherence, respect physical plausibility, and retain visual appeal. The lack of robust automated edit-quality metrics hinders reliable automation at scale. We present an automated, modular pipeline that mines high-fidelity triplets across domains, resolutions, instruction complexities, and styles. Built on public generative models and running without human intervention, our system uses a task-tuned Gemini validator to score instruction adherence and aesthetics directly, removing any need for segmentation or grounding models. Inversion and compositional bootstrapping enlarge the mined set by approx. 2.6x, enabling large-scale high-fidelity training data. By automating the most repetitive annotation steps, the approach allows a new scale of training without human labeling effort. To democratize research in this resource-intensive area, we release NHR-Edit, an open dataset of 720k high-quality triplets, curated at industrial scale via millions of guided generations and validator passes, and we analyze the pipeline's stage-wise survival rates, providing a framework for estimating computational effort across different model stacks. In the largest cross-dataset evaluation, it surpasses all public alternatives. We also release Bagel-NHR-Edit, a fine-tuned Bagel model with state-of-the-art metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。