构建首个统一图文生成评估基准,精准衡量模型复杂多模态能力。
UEval: A Benchmark for Unified Multimodal Generation
- 设计基于评分标准的自动化评估体系,提升评测细粒度。
- 1000个真实任务问题,涵盖多类推理类型,测试综合生成能力。
- 发现推理能力显著提升生成质量,适合研究多模态模型的开发者。
我们提出UEval,一个用于评估统一多模态生成模型(可同时生成图像和文本)的基准。该基准包含1000个专家精心设计的问题,源自8个真实应用场景,要求模型输出图文结合的内容。问题覆盖从分步指南到教科书解释等多样推理类型。传统基于LLM的评估方法难以捕捉细节,因此我们采用基于评分标准的系统:先用多模态大模型生成初始评分维度,再由人工专家修正验证。最终共获得10,417条经验证的评分标准,支持可扩展的细粒度自动评分。当前统一模型表现有限:GPT-5-Thinking得分为66.4/100,最佳开源模型仅49.1。实验表明,具备推理能力的模型表现更优,且将推理轨迹迁移至非推理模型可显著缩小性能差距,说明推理对复杂多模态理解与生成至关重要。
原文摘要 · Abstract (English)
We introduce UEval, a benchmark to evaluate unified models, i.e., models capable of generating both images and text. UEval comprises 1,000 expert-curated questions that require both images and text in the model output, sourced from 8 real-world tasks. Our curated questions cover a wide range of reasoning types, from step-by-step guides to textbook explanations. Evaluating open-ended multimodal generation is non-trivial, as simple LLM-as-a-judge methods can miss the subtleties. Different from previous works that rely on multimodal Large Language Models (MLLMs) to rate image quality or text accuracy, we design a rubric-based scoring system in UEval. For each question, reference images and text answers are provided to a MLLM to generate an initial rubric, consisting of multiple evaluation criteria, and human experts then refine and validate these rubrics. In total, UEval contains 10,417 validated rubric criteria, enabling scalable and fine-grained automatic scoring. UEval is challenging for current unified models: GPT-5-Thinking scores only 66.4 out of 100, while the best open-source model reaches merely 49.1. We observe that reasoning models often outperform non-reasoning ones, and transferring reasoning traces from a reasoning model to a non-reasoning model significantly narrows the gap. This suggests that reasoning may be important for tasks requiring complex multimodal understanding and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。