评测文生图模型的推理能力,发现现有模型表现普遍不足。
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
- 构建涵盖多种推理类型的图文生成评测集
- 提出三维度评分体系,精准衡量推理准确性
- 16个模型测试显示推理能力严重滞后,适合研究者参考
现实场景中的文生图任务常需推理能力,例如生成‘被咬过的苹果在空气中放置一周以上’需理解时间衰变与常识概念。尽管当前文生图模型已能生成逼真图像,其推理能力仍薄弱且缺乏有效评估。为此,我们提出R2I-Bench,一个专门用于严谨评估推理驱动文生图的综合性基准。该基准包含精心设计的数据实例,覆盖常识、数学、逻辑、组合、数值、因果及概念混合等核心推理类别。为实现细粒度评估,我们设计了R2IScore,一种基于实例定制的问答式评分指标,从文本-图像对齐、推理准确性和图像质量三个维度进行评价。对16个代表性文生图模型(包括使用先进语言与图像模型解耦推理与生成的强基线框架)的广泛实验表明,各模型推理性能普遍有限,凸显下一代文生图系统亟需更鲁棒的推理感知架构。
原文摘要 · Abstract (English)
Reasoning is a fundamental capability often required in real-world text-to-image (T2I) generation, e.g., generating ``a bitten apple that has been left in the air for more than a week`` necessitates understanding temporal decay and commonsense concepts. While recent T2I models have made impressive progress in producing photorealistic images, their reasoning capability remains underdeveloped and insufficiently evaluated. To bridge this gap, we introduce R2I-Bench, a comprehensive benchmark specifically designed to rigorously assess reasoning-driven T2I generation. R2I-Bench comprises meticulously curated data instances, spanning core reasoning categories, including commonsense, mathematical, logical, compositional, numerical, causal, and concept mixing. To facilitate fine-grained evaluation, we design R2IScore, a QA-style metric based on instance-specific, reasoning-oriented evaluation questions that assess three critical dimensions: text-image alignment, reasoning accuracy, and image quality. Extensive experiments with 16 representative T2I models, including a strong pipeline-based framework that decouples reasoning and generation using the state-of-the-art language and image generation models, demonstrate consistently limited reasoning performance, highlighting the need for more robust, reasoning-aware architectures in the next generation of T2I systems. Project Page: https://r2i-bench.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。