arXiv:2502.07870cs.CV2025-02被引 23

构建500万张长文本图像数据集,推动文本生成图像模型发展

TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation

  • 构建500万张含长文本的生成图像数据集,覆盖多种场景
  • 在TextAtlasEval测试集上,顶级模型仍面临显著挑战
  • 适合研究长文本生成、图文对齐与评估基准的学者使用

近年来,文本条件图像生成受到广泛关注,生成文本提示的长度和复杂度持续增加。日常生活中,广告、信息图和标识等场景普遍存在密集复杂的文字内容,图文融合是传递复杂信息的关键。然而,长文本图像生成仍是重大挑战,主要受限于现有数据集多聚焦短文本。为此,我们提出TextAtlas5M,一个专为评估长文本渲染能力而设计的大规模数据集,包含500万张涵盖多种数据类型的生成与收集图像,支持大模型在长文本图像生成任务上的全面评估。我们进一步构建了3000张人工优化的测试集TextAtlasEval,覆盖三个数据领域,建立当前最广泛的文本条件生成基准。评估表明,即使最先进的专有模型(如GPT4o with DallE-3)在TextAtlasEval上仍表现困难,开源模型差距更大。这证明TextAtlas5M是未来文本条件图像生成模型训练与评估的重要资源。

原文摘要 · Abstract (English)

Text-conditioned image generation has gained significant attention in recent years and are processing increasingly longer and comprehensive text prompt. In everyday life, dense and intricate text appears in contexts like advertisements, infographics, and signage, where the integration of both text and visuals is essential for conveying complex information. However, despite these advances, the generation of images containing long-form text remains a persistent challenge, largely due to the limitations of existing datasets, which often focus on shorter and simpler text. To address this gap, we introduce TextAtlas5M, a novel dataset specifically designed to evaluate long-text rendering in text-conditioned image generation. Our dataset consists of 5 million long-text generated and collected images across diverse data types, enabling comprehensive evaluation of large-scale generative models on long-text image generation. We further curate 3000 human-improved test set TextAtlasEval across 3 data domains, establishing one of the most extensive benchmarks for text-conditioned generation. Evaluations suggest that the TextAtlasEval benchmarks present significant challenges even for the most advanced proprietary models (e.g. GPT4o with DallE-3), while their open-source counterparts show an even larger performance gap. These evidences position TextAtlas5M as a valuable dataset for training and evaluating future-generation text-conditioned image generation models.

文本生成图像长文本数据集图文对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。