arXiv:2608.18076cs.CVcs.AI2026-08被引 1

按能力演化顺序设计数据,让图像生成模型更全面、更连贯。

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

论文配图:From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
图 1 · 摘自论文原文
  • 以生成能力为驱动,构建分阶段的数据流水线
  • 训练出30亿和60亿参数的多模态扩散模型
  • 适合需要强泛化与跨任务迁移的图像生成研究者

大规模图像生成虽受益于数据规模、质量与重标注提升,但传统流程通常孤立优化特定任务数据集。核心挑战不仅在于如何构建每个任务的数据集,更在于如何根据生成能力间的依赖关系组织异构监督。本文提出一种能力驱动的数据基础设施,将能力特定的监督构建与能力对齐的课程调度相结合。三个专用但可互操作的数据引擎分别构建文本-图像对齐、图像间转换、图像-知识关联的互补关系监督,同时通过标题专家在任务间与粒度上对齐文本到图像(T2I)与编辑监督。多阶段课程沿能力获取的依赖顺序,联合演化任务构成、视觉概念分布、数据质量与图像分辨率,并通过能力感知评估闭环反馈:针对性检索、专家构建与差距感知重采样。在规模上,该框架构建了4.4亿张图像的T2I语料库、1.2亿对编辑样本及超过2700万张图像-实体对。基于此基础设施,我们从零开始训练了30亿与60亿参数的多模态扩散模型。在CPI-Bench上进行量化评估,并在多种文本到图像与编辑场景中开展定性分析。实验结果表明,模型具备广泛视觉覆盖、多样渲染能力以及有效的生成能力迁移。

原文摘要 · Abstract (English)

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf{capability-driven data infrastructure} that couples capability-specific supervision construction with capability-aligned curriculum scheduling. Its three specialized yet interoperable data engines build complementary relational supervision for text-image grounding, inter-image transformation, and image-knowledge association, while caption experts align T2I and editing supervision across tasks and granularities. A multi-stage curriculum jointly evolves task composition, visual-concept distribution, data quality, and image resolution along the dependency order of capability acquisition, with capability-aware evaluation closing the loop through targeted retrieval, expert construction, and gap-aware resampling. At scale, the framework curates a 440M-image T2I corpus, 120M editing pairs, and over 27M image-entity pairs. With this infrastructure, we train multimodal diffusion models at two scales from scratch, with 3B and 6B sizes respectively. We conduct quantitative evaluation on CPI-Bench, along with qualitative evaluations across diverse text-to-image and editing scenarios. Experimental results present broad visual coverage, versatile rendering, and effective transfer across generative capabilities.

图像生成扩散模型数据工程能力演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。