首个支持跨任务与多维度评估的统一评测框架,解决大模型生成质量评估难题。
FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization
- 构建包含112个细粒度维度的层级评估体系
- 基于60.4万样本数据集实现对未见维度的泛化评估
- 适合需要全面评估多模态生成质量的研究者使用
多模态大模型的开放生成输出评估已成为瓶颈,现有基于模型自评的方法局限于特定任务和维度。本文提出统一细粒度评估框架UFEval,涵盖自然语言生成、图像理解、图像生成及图文混合生成四类任务。为支撑训练,构建了包含60.4万个成对样本、32.5万条标注的FRABench数据集,基于112个细粒度维度的层级分类体系,融合人工与GPT-4o标注。实验表明,针对特定维度的学习可泛化至未见维度,联合学习多种任务与维度能产生显著协同增益。
原文摘要 · Abstract (English)
Evaluating open-ended outputs of Multimodal Large Language Models has become a bottleneck as model capabilities, task diversity, and modality rapidly expand. Existing ``MLLM-as-a-Judge'' evaluators, though promising, remain constrained to specific tasks and aspects. In this paper, we argue that, on one hand, based on the interconnected nature of aspects, learning specific aspects can generalize to unseen aspects; on the other hand, jointly learning to assess multiple visual aspects and tasks may foster a synergistic effect. To this end, we propose UFEval, the first unified fine-grained evaluator with task and aspect generalization for four evaluation tasks -- Natural Language Generation, Image Understanding, Image Generation, and Interleaved Text-and-Image Generation. However, training such a unified evaluator is hindered by the lack of a large-scale, multi-modal, and aspect-level resource. To address this gap, we introduce FRABench, a comprehensive fine-grained evaluation dataset. Specifically, (1) We first construct a hierarchical aspect taxonomy encompassing 112 distinct aspects across the aforementioned four tasks. (2) Based on this taxonomy, we create FRABench, comprising 60.4k pairwise samples with 325k evaluation labels obtained from a combination of human and GPT-4o annotations. (3) Finally, leveraging FRABench, we develop UFEval, a unified fine-grained evaluator. Experiments show that learning on specific aspects enables UFEval to generalize to unseen aspects, and joint learning to assess diverse visual tasks and aspects can lead to substantial mutual benefits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。