构建多语言、细粒度的图文生成评估基准,全面检验模型语义理解能力。
UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- 设计600个分层提示,覆盖5大主题20子类,支持中英文长短文本
- 涵盖10个主维度27个子维度,每个提示测试多个语义点
- 基于Gemini-2.5-Pro构建评估流水线,可离线使用训练模型
近期图文生成进展凸显可靠评估基准的重要性,但现有基准存在提示场景单一、缺乏多语言支持、评估维度粗略等问题。为此,我们提出UniGenBench++,一个统一的语义评估基准。该基准包含600个分层组织的提示,覆盖5大主题和20个子主题,确保真实场景多样性;在10个主维度和27个子维度上评估模型语义一致性,每个提示对应多个测试点。为评估模型对语言和提示长度的鲁棒性,提供中英文短/长版本提示。利用闭源多模态大模型Gemini-2.5-Pro的通用知识与细粒度图像理解能力,构建可靠评估流程,并训练出可离线使用的评估模型。通过全面评测开源与闭源模型,系统揭示其在各类任务中的优劣表现。
原文摘要 · Abstract (English)
Recent progress in text-to-image (T2I) generation underscores the importance of reliable benchmarks in evaluating how accurately generated images reflect the semantics of their textual prompt. However, (1) existing benchmarks lack the diversity of prompt scenarios and multilingual support, both essential for real-world applicability; (2) they offer only coarse evaluations across primary dimensions, covering a narrow range of sub-dimensions, and fall short in fine-grained sub-dimension assessment. To address these limitations, we introduce UniGenBench++, a unified semantic assessment benchmark for T2I generation. Specifically, it comprises 600 prompts organized hierarchically to ensure both coverage and efficiency: (1) spans across diverse real-world scenarios, i.e., 5 main prompt themes and 20 subthemes; (2) comprehensively probes T2I models' semantic consistency over 10 primary and 27 sub evaluation criteria, with each prompt assessing multiple testpoints. To rigorously assess model robustness to variations in language and prompt length, we provide both English and Chinese versions of each prompt in short and long forms. Leveraging the general world knowledge and fine-grained image understanding capabilities of a closed-source Multi-modal Large Language Model (MLLM), i.e., Gemini-2.5-Pro, an effective pipeline is developed for reliable benchmark construction and streamlined model assessment. Moreover, to further facilitate community use, we train a robust evaluation model that enables offline assessment of T2I model outputs. Through comprehensive benchmarking of both open- and closed-sourced T2I models, we systematically reveal their strengths and weaknesses across various aspects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。