首个统一多模态模型评估框架,无需额外数据与标注。
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
- 构建无外部依赖的统一评估框架,支持理解与生成任务
- 新基准UniBench含81个细粒度标签,挑战性更强
- 评分指标与人工评价高度一致,适合研究者快速对比模型
统一多模态理解和生成模型因具备更强指令遵循能力且减少模型冗余而迅速受到关注。然而,当前缺乏适用于此类模型的统一评估框架,现有方法依赖多个任务特定基准,存在整体结果缺失、依赖额外评估模型、需大量带标签图像、基准多样性不足以及评估指标对指令遵循能力覆盖有限等问题。为此,我们提出UniEval,首个无需额外模型、图像或标注的统一多模态模型评估框架。该框架包含综合性基准UniBench(支持统一与视觉生成模型)及对应的UniScore评估指标。UniBench包含81个细粒度标签,实现高多样性。实验表明,UniBench比现有基准更具挑战性,UniScore与人工评价高度一致,超越现有指标。我们还对当前主流统一与视觉生成模型进行了全面评估,揭示了统一模型的独特价值。
原文摘要 · Abstract (English)
The emergence of unified multimodal understanding and generation models is rapidly attracting attention because of their ability to enhance instruction-following capabilities while minimizing model redundancy. However, there is a lack of a unified evaluation framework for these models, which would enable an elegant, simplified, and overall evaluation. Current models conduct evaluations on multiple task-specific benchmarks, but there are significant limitations, such as the lack of overall results, errors from extra evaluation models, reliance on extensive labeled images, benchmarks that lack diversity, and metrics with limited capacity for instruction-following evaluation. To tackle these challenges, we introduce UniEval, the first evaluation framework designed for unified multimodal models without extra models, images, or annotations. This facilitates a simplified and unified evaluation process. The UniEval framework contains a holistic benchmark, UniBench (supports both unified and visual generation models), along with the corresponding UniScore metric. UniBench includes 81 fine-grained tags contributing to high diversity. Experimental results indicate that UniBench is more challenging than existing benchmarks, and UniScore aligns closely with human evaluations, surpassing current metrics. Moreover, we extensively evaluated SoTA unified and visual generation models, uncovering new insights into Univeral's unique values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。