arXiv:2411.15296cs.CVcs.AI2024-11综述被引 82

系统梳理多模态大模型评估方法,助研究者高效评测模型性能。

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

  • 按能力类型分类评估基准,涵盖基础能力、自省分析与应用扩展。
  • 总结基准构建全流程:数据收集、标注规范与注意事项。
  • 提出评估体系框架,含评判标准、度量指标与工具链,适合科研参考。

作为人工智能通用化的重要方向,多模态大语言模型(MLLMs)受到产业界与学术界的广泛关注。基于预训练大语言模型,该类模型进一步发展出强大的多模态感知与推理能力,例如根据流程图生成代码或基于图像创作故事。在模型研发过程中,评估至关重要,能提供直观反馈以指导优化。不同于传统单一任务(如图像分类)的训练-评估-测试范式,MLLMs的多功能性催生了多种新型评估基准与方法。本文旨在全面综述多模态大模型的评估体系,涵盖四个核心方面:1)按评估能力划分的基准类型,包括基础能力、模型自省分析与扩展应用;2)基准构建的典型流程,包含数据采集、标注及注意事项;3)系统化的评估方式,涵盖评判机制、度量指标与工具集;4)未来评估基准的发展展望。本工作旨在帮助研究人员根据需求高效评估MLLMs,并激发更优评估方法的设计,推动该领域持续发展。

原文摘要 · Abstract (English)

As a prominent direction of Artificial General Intelligence (AGI), Multimodal Large Language Models (MLLMs) have garnered increased attention from both industry and academia. Building upon pre-trained LLMs, this family of models further develops multimodal perception and reasoning capabilities that are impressive, such as writing code given a flow chart or creating stories based on an image. In the development process, evaluation is critical since it provides intuitive feedback and guidance on improving models. Distinct from the traditional train-eval-test paradigm that only favors a single task like image classification, the versatility of MLLMs has spurred the rise of various new benchmarks and evaluation methods. In this paper, we aim to present a comprehensive survey of MLLM evaluation, discussing four key aspects: 1) the summarised benchmarks types divided by the evaluation capabilities, including foundation capabilities, model self-analysis, and extented applications; 2) the typical process of benchmark counstruction, consisting of data collection, annotation, and precautions; 3) the systematic evaluation manner composed of judge, metric, and toolkit; 4) the outlook for the next benchmark. This work aims to offer researchers an easy grasp of how to effectively evaluate MLLMs according to different needs and to inspire better evaluation methods, thereby driving the progress of MLLM research.

多模态大模型评估综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。