arXiv:2506.00883cs.CL2025-06被引 3

用少样本问答快速评估多模态大模型性能

Improve MLLM Benchmark Efficiency through Interview

  • 基于模型表现给题目打难易标签,筛选关键问题进行测试
  • 仅需少量问题即可准确预测模型在全数据集的表现
  • 适合需要高效评测的科研人员和模型开发者

多模态大语言模型(MLLM)快速发展,催生了大量评估数据集。然而,在大规模数据上进行全面问答测试耗时耗力。为此,本文提出一种名为MLLM Interview(MITV)的策略,通过少量提问快速获取模型性能指标。首先,基于现有评估数据集,结合典型MLLM在其中的表现,为题目添加难度标签,构建访谈数据集;其次,提出一种渐进式测试策略:先用少量主题初探模型能力,再持续挑战其极限。大量实验表明,该方法在多个MLLM基准数据集上表现优异,仅需少量问答即可高效获得模型评估结果。

原文摘要 · Abstract (English)

The rapid development of Multimodal Large Language Models (MLLM) has led to a wide range of MLLM applications, and a number of benchmark datasets have sprung up in order to assess MLLM abilities. However, full-coverage Q&A testing on large-scale data is resource-intensive and time-consuming. To address this issue, we propose the MLLM Interview (MITV) strategy, which aims to quickly obtain MLLM performance metrics by quizzing fewer question. First, First, we constructed the interview dataset, which was built on an existing MLLM assessment dataset, by adding difficulty labels based on the performance of some typical MLLMs in this dataset. Second, we propose an MLLM Interview strategy, which obtains an initial performance situation of the large model by quizzing a small number of topics and then continuously tries to test the model's limits. Through extensive experiments, the result shows that the MITV strategy proposed in this paper performs well on MLLM benchmark datasets, and it is able to obtain the model evaluation capability faster through a small number of questions and answers.

多模态模型评测效率小样本评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。