arXiv:2409.18142cs.AIcs.MM2024-09综述被引 31

梳理211个多模态大模型评测基准,厘清评估标准与数据设计。

A Survey on Multimodal Benchmarks: In the Era of Large AI Models

  • 系统分析211个基准,覆盖理解、推理、生成和应用四类任务
  • 揭示不同模态下任务设计与评价指标的差异与共性
  • 适合研究者快速了解多模态模型评测现状

多模态大语言模型(MLLMs)的快速发展推动了人工智能在多模态内容理解与生成能力上的显著提升。尽管以往研究主要关注模型架构与训练方法,但对评估这些模型所用基准的系统性分析仍显不足。本综述系统调研了211个涵盖理解、推理、生成和应用四大核心领域的多模态基准,深入分析了任务设计、评价指标与数据集构建方式,覆盖多种模态。本文旨在为多模态大模型研究提供全面的基准评估视角,并指明未来研究方向。相关论文资源已收录于配套GitHub仓库。

原文摘要 · Abstract (English)

The rapid evolution of Multimodal Large Language Models (MLLMs) has brought substantial advancements in artificial intelligence, significantly enhancing the capability to understand and generate multimodal content. While prior studies have largely concentrated on model architectures and training methodologies, a thorough analysis of the benchmarks used for evaluating these models remains underexplored. This survey addresses this gap by systematically reviewing 211 benchmarks that assess MLLMs across four core domains: understanding, reasoning, generation, and application. We provide a detailed analysis of task designs, evaluation metrics, and dataset constructions, across diverse modalities. We hope that this survey will contribute to the ongoing advancement of MLLM research by offering a comprehensive overview of benchmarking practices and identifying promising directions for future work. An associated GitHub repository collecting the latest papers is available.

多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。