arXiv:2606.20100cs.CV2026-06

为文生图模型设计多维度诊断基准,精准定位生成缺陷

WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization

论文配图:WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization
图 1 · 摘自论文原文
  • 构建4000条中英平衡的测试提示,分多维标签细化任务
  • 引入视觉语言模型评估新指标,从三方面衡量生成质量
  • 不仅给出结果,还提供推理过程,可验证评估可靠性

近期文生图模型在仅凭文本输入生成高度逼真图像方面表现出色。尽管现有基准可在一定程度上评估模型生成能力,但难以全面准确地衡量多维度性能,常无法揭示模型在特定类别中的固有缺陷。为此,我们提出WeGenBench,一个旨在全面、多角度评估文生图能力的新基准。该基准包含4000条测试提示,分为两大类,中英文均衡,用于评估双语及跨文化生成能力。除宏观场景分类外,每条提示均标注多维标签,针对不同语言的内容与挑战进行细化。通过融合场景分类与多维标签的跨维度评估机制,WeGenBench能精准定位模型在特定生成类别中的不足。此外,为更准确衡量生成质量,我们设计并验证了多个基于视觉语言模型(VLMs)的新评估指标,从三个核心方面评估模型在领域特定任务上的表现。关键在于,该方法同时产出评估结果与详细推理轨迹,支持对评估结果准确性和合理性进行严格验证。最后,我们在当前最先进方法上进行了系统性基准测试,并深入分析了现有模型存在的局限。

原文摘要 · Abstract (English)

Recent text-to-image generation models have demonstrated remarkable capabilities in synthesizing highly realistic images from text inputs alone. Although existing benchmarks can evaluate the generation capabilities of various models to some extent, they struggle to comprehensively and accurately measure performance across multiple dimensions, often failing to reveal the inherent deficiencies of models in specific categories. To address these limitations, we propose WeGenBench, a novel benchmark designed for the comprehensive, multi-perspective evaluation of text-to-image generation capabilities. Our benchmark comprises a total of 4,000 test prompts across two primary categories, meticulously balanced between Chinese and English to evaluate bilingual and cross-cultural generation capabilities. Beyond macroscopic scene classification, we annotate each prompt with multi-dimensional tags tailored to the distinct content and challenges of each language, thereby refining the generation tasks into more specific sub-categories. Through a cross-dimensional evaluation mechanism leveraging both scene classifications and multi-dimensional tags, WeGenBench can precisely pinpoint model shortcomings in specific generation categories. Furthermore, to measure generation quality more accurately, we design and validate several novel evaluation metrics by integrating Vision-Language Models (VLMs), which assess model performance on domain-specific tasks from three core aspects. Crucially, our approach yields both the assessment outcomes and the detailed reasoning trajectories, facilitating a rigorous verification of the accuracy and soundness of the evaluation results. Finally, we conduct systematic benchmarking on current state-of-the-art methods and provide an in-depth analysis of the limitations present in existing models.

文生图评估基准多维度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。