arXiv:2509.14886cs.CLcs.AI2025-09

模仿面试流程,用更少问题高效评估多模态大模型性能

A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation

  • 设计两阶段面试机制,动态调整提问权重
  • 相比随机采样,相关性提升最高达17.6%(PLCC)
  • 适合大规模模型评测,兼顾效率与可靠性

多模态大语言模型(MLLMs)的快速发展催生了大量评测基准。然而,传统全覆盖问答评测存在高度冗余和效率低下问题。受人类面试流程启发,我们提出一种多对一面试范式,用于高效评估MLLMs。该框架包含三部分:(i) 两阶段面试策略,包括预面试与正式面试阶段;(ii) 动态调整面试官权重以保障公平性;(iii) 自适应选择问题难度层级。在多个基准上的实验表明,该范式相较于随机采样,在全覆盖结果上显著提升相关性,PLCC最高提升17.6%,SRCC提升16.7%,同时大幅减少所需问题数量。结果证明,该范式为大规模MLLM评测提供了可靠且高效的替代方案。

原文摘要 · Abstract (English)

The rapid progress of Multi-Modal Large Language Models (MLLMs) has spurred the creation of numerous benchmarks. However, conventional full-coverage Question-Answering evaluations suffer from high redundancy and low efficiency. Inspired by human interview processes, we propose a multi-to-one interview paradigm for efficient MLLM evaluation. Our framework consists of (i) a two-stage interview strategy with pre-interview and formal interview phases, (ii) dynamic adjustment of interviewer weights to ensure fairness, and (iii) an adaptive mechanism for question difficulty-level chosen. Experiments on different benchmarks show that the proposed paradigm achieves significantly higher correlation with full-coverage results than random sampling, with improvements of up to 17.6% in PLCC and 16.7% in SRCC, while reducing the number of required questions. These findings demonstrate that the proposed paradigm provides a reliable and efficient alternative for large-scale MLLM benchmarking.

多模态模型评测高效评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。