arXiv:2509.09680cs.CVcs.CL2025-09被引 33

构建600万张图文推理数据集,推动开源文生图模型发展

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark

  • 构建600万张图像与2000万条中英文描述的推理数据集
  • 首次在长文本生成任务中实现基于思维链的精准评估
  • 适合研究文生图推理能力与模型评测的研究者使用

开源文生图模型的发展受限于缺乏大规模、聚焦推理的数据集与全面的评估基准,导致性能落后于闭源系统。为此,我们推出FLUX-Reason-6M与PRISM-Bench(精确且鲁棒的图像生成评估基准)。FLUX-Reason-6M包含600万张高质量由FLUX生成的图像及2000万条中英文双语描述,专为复杂推理训练设计。图像按想象、实体、文字渲染、风格、情感、构图六大特征组织,并引入显式生成思维链(GCoT)以分解生成步骤。整个数据构建耗时15,000个A100 GPU日,为社区提供此前仅限大型工业实验室才能获取的资源。PRISM-Bench提出七项新评估赛道,包括采用GCoT的长文本挑战。通过精心设计的提示,结合先进视觉语言模型,实现对提示-图像匹配度与图像美学的人类对齐评估。我们在PRISM-Bench上对19个主流模型进行评估,揭示了显著性能差距并指明改进方向。相关数据集、基准与代码均已开源,以推动下一代推理导向的文生图生成技术发展。

原文摘要 · Abstract (English)

The advancement of open-source text-to-image (T2I) models has been hindered by the absence of large-scale, reasoning-focused datasets and comprehensive evaluation benchmarks, resulting in a performance gap compared to leading closed-source systems. To address this challenge, We introduce FLUX-Reason-6M and PRISM-Bench (Precise and Robust Image Synthesis Measurement Benchmark). FLUX-Reason-6M is a massive dataset consisting of 6 million high-quality FLUX-generated images and 20 million bilingual (English and Chinese) descriptions specifically designed to teach complex reasoning. The image are organized according to six key characteristics: Imagination, Entity, Text rendering, Style, Affection, and Composition, and design explicit Generation Chain-of-Thought (GCoT) to provide detailed breakdowns of image generation steps. The whole data curation takes 15,000 A100 GPU days, providing the community with a resource previously unattainable outside of large industrial labs. PRISM-Bench offers a novel evaluation standard with seven distinct tracks, including a formidable Long Text challenge using GCoT. Through carefully designed prompts, it utilizes advanced vision-language models for nuanced human-aligned assessment of prompt-image alignment and image aesthetics. Our extensive evaluation of 19 leading models on PRISM-Bench reveals critical performance gaps and highlights specific areas requiring improvement. Our dataset, benchmark, and evaluation code are released to catalyze the next wave of reasoning-oriented T2I generation. Project page: https://flux-reason-6m.github.io/ .

文生图推理数据集模型评测多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。