arXiv:2605.21272cs.CVcs.AI2026-05

构建了1亿+规模的高质量图文数据集,助力文本生成图像研究开源可复现。

MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

论文配图:MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset
图 1 · 摘自论文原文
  • 从29亿原始配对中筛选并去重,经多轮过滤与多模型重描述
  • 包含约1.049亿张图像-文本对,支持短长描述与预计算嵌入
  • 适合作为大模型训练基础数据集,尤其适合追求可复现性的研究者

训练大型文本到图像模型需要高质量、经过精心筛选的数据集,涵盖多样化内容和详细描述。然而,在大规模下收集、过滤、去重和重新标注语料库的成本与复杂性,阻碍了该领域的开放与可复现研究。我们提出 MONET,一个开放的 Apache 2.0 许可数据集,包含约 1.049 亿个图像-文本对,源自 29 亿条原始配对,通过多个阶段的安全过滤、基于领域过滤、精确与近似重复去除,以及使用多种视觉-语言模型进行短至长形式的重描述,并进一步通过合成样本增强。每张图像均附带预计算嵌入与标注,以加速下游应用。为验证 MONET 的有效性,我们在其上仅训练了一个 40 亿参数的潜在扩散模型,达到与现有方法相当的 GenEval 与 DPG 得分,证明该数据集降低了大规模、可复现文本到图像研究的门槛。

原文摘要 · Abstract (English)

Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of collecting, filtering, deduplicating, and re-captioning such corpora at scale hinders open and reproducible research in the field. We introduce MONET, an open Apache 2.0 dataset of approx. 104.9M image--text pairs collected from 2.9B raw pairs across heterogeneous open sources through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with multiple vision-language models covering short to long-form descriptions, and further augmented with synthetically generated samples. Each image is shipped with pre-computed embeddings and annotations to accelerate downstream use. To validate the effectiveness of MONET, we train a 4B-parameter latent diffusion model exclusively on it and reach competitive GenEval and DPG scores, demonstrating that our dataset lowers the barrier to large-scale, reproducible text-to-image research.

图文数据集文本生成图像开源数据扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。