arXiv:2410.10563cs.CV2024-10被引 41

构建超500项真实多模态任务评测集,支持多样化输出格式。

MEGA-Bench: Scaling Multimodal Evaluation to over 500 Real-World Tasks

  • 收集505个真实任务、8000+样本,覆盖多种输入输出形式。
  • 设计40+评估指标,支持数值、代码、JSON等10余种输出格式。
  • 提供细粒度能力报告,适合模型性能对比与应用选型。

我们提出MEGA-Bench,一个涵盖超过500项真实世界多模态任务的评测基准,以应对终端用户高度异构的应用场景。目标是构建高质量、多样化的数据样本,全面覆盖多模态任务空间,同时实现低成本、高精度的模型评估。具体地,我们从16位专家标注者处收集了505个真实任务,包含超过8,000个样本,广泛覆盖多模态任务类型。不同于将问题统一为标准选择题(如MMMU、MMBench、MMT-Bench),MEGA-Bench支持数字、短语、代码、LaTeX、坐标、JSON、自由文本等多种输出格式。为此,我们开发了40多个评估指标以适配不同输出形式。相比现有基准,MEGA-Bench可在应用类别、输入类型、输出格式、技能维度等多个层面提供细粒度的能力报告,支持用户深度交互与可视化分析。我们在MEGA-Bench上对多种前沿视觉-语言模型进行了评估,以揭示其在各维度上的综合表现。

原文摘要 · Abstract (English)

We present MEGA-Bench, an evaluation suite that scales multimodal evaluation to over 500 real-world tasks, to address the highly heterogeneous daily use cases of end users. Our objective is to optimize for a set of high-quality data samples that cover a highly diverse and rich set of multimodal tasks, while enabling cost-effective and accurate model evaluation. In particular, we collected 505 realistic tasks encompassing over 8,000 samples from 16 expert annotators to extensively cover the multimodal task space. Instead of unifying these problems into standard multi-choice questions (like MMMU, MMBench, and MMT-Bench), we embrace a wide range of output formats like numbers, phrases, code, \LaTeX, coordinates, JSON, free-form, etc. To accommodate these formats, we developed over 40 metrics to evaluate these tasks. Unlike existing benchmarks, MEGA-Bench offers a fine-grained capability report across multiple dimensions (e.g., application, input type, output format, skill), allowing users to interact with and visualize model capabilities in depth. We evaluate a wide variety of frontier vision-language models on MEGA-Bench to understand their capabilities across these dimensions.

多模态评测真实任务能力分析输出泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。