arXiv:2501.13953cs.CLcs.AI2025-01ACL被引 20

分析多模态大模型评测的冗余问题,提出构建高效评测的三条原则。

Redundancy Principles for MLLMs Benchmarks

  • 从能力维度、题目数量、领域重叠三方面识别评测冗余
  • 在20多个基准上分析数百个模型表现,量化冗余程度
  • 适合评测设计者、研究者参考,避免重复建设

随着多模态大语言模型(MLLMs)的快速迭代和领域需求演变,每年产生的评测基准已超过百个,导致显著冗余。本文从三个关键视角分析冗余:1)评测能力维度的重复;2)测试题目的数量冗余;3)特定领域的跨基准重叠。通过对超过20个基准上数百个MLLMs的性能进行综合分析,我们定量评估了现有评测中的冗余水平,为未来构建高效、少冗余的评测体系提供洞察与策略支持。代码已开源:https://github.com/zzc-1998/Benchmark-Redundancy。

原文摘要 · Abstract (English)

With the rapid iteration of Multi-modality Large Language Models (MLLMs) and the evolving demands of the field, the number of benchmarks produced annually has surged into the hundreds. The rapid growth has inevitably led to significant redundancy among benchmarks. Therefore, it is crucial to take a step back and critically assess the current state of redundancy and propose targeted principles for constructing effective MLLM benchmarks. In this paper, we focus on redundancy from three key perspectives: 1) Redundancy of benchmark capability dimensions, 2) Redundancy in the number of test questions, and 3) Cross-benchmark redundancy within specific domains. Through the comprehensive analysis over hundreds of MLLMs' performance across more than 20 benchmarks, we aim to quantitatively measure the level of redundancy lies in existing MLLM evaluations, provide valuable insights to guide the future development of MLLM benchmarks, and offer strategies to refine and address redundancy issues effectively. The code is available at https://github.com/zzc-1998/Benchmark-Redundancy.

评测基准冗余分析MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。