arXiv:2508.17472cs.CV2025-08被引 34
评测文生图模型的推理能力,涵盖四类复杂语义理解任务。
T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation

- 构建四维评测体系:隐喻理解、文本图像设计、实体推理、科学推理。
- 提出两阶段评估流程,兼顾推理准确率与图像质量。
- 首次系统评测主流文生图模型在复杂推理任务中的表现。
我们提出了T2I-ReasonBench,一个用于评估文本到图像(T2I)模型推理能力的基准测试。该基准包含四个维度:隐喻理解、文本图像设计、实体推理和科学推理。我们设计了一个两阶段评估协议,以衡量模型的推理准确性和生成图像质量。通过对多种T2I生成模型进行基准测试,提供了对其性能的全面分析。
原文摘要 · Abstract (English)
We propose T2I-ReasonBench, a benchmark evaluating reasoning capabilities of text-to-image (T2I) models. It consists of four dimensions: Idiom Interpretation, Textual Image Design, Entity-Reasoning and Scientific-Reasoning. We propose a two-stage evaluation protocol to assess the reasoning accuracy and image quality. We benchmark various T2I generation models, and provide comprehensive analysis on their performances.
文生图推理能力基准测试
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。