arXiv:2505.21389cs.CL2025-05ACL被引 3

用智能代理动态选题,让多模态模型评测成本降为原来的4%。

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs

  • 基于项目反应理论与自适应代理,实时选择最有信息量的题目。
  • 仅用4%数据就达到90%以上全量评估的排名准确率。
  • 适合需要高效评测多模态大模型的研究者和工程师。

评估多模态大语言模型(MLLMs)的成本正日益攀升,因基准测试规模扩大且跨模态复杂性增加,导致评分工作量巨大。为此,我们提出AutoJudger,一种由智能体驱动的高效自适应评测框架。该框架采用项目反应理论(IRT)估计题目难度,并通过自主评估代理根据模型实时表现动态选择最具信息量的测试题。其核心包含两个关键组件:语义感知的检索机制,确保所选题目覆盖视觉与语言模态中的多样且具有挑战性的场景;以及动态记忆模块,持续记录过往题目的上下文统计信息,以实现全局连贯的题目选择。在四个代表性多模态基准上的大量实验表明,该自适应框架显著降低评测开销——在MMT-Bench上仅需4%的数据即可实现超过90%的排名准确率,媲美全量评估效果。

原文摘要 · Abstract (English)

Evaluating multimodal large language models (MLLMs) is increasingly expensive, as the growing size and cross-modality complexity of benchmarks demand significant scoring efforts. To tackle with this difficulty, we introduce AutoJudger, an agent-driven framework for efficient and adaptive benchmarking of MLLMs that tackles this escalating cost. AutoJudger employs the Item Response Theory (IRT) to estimate the question difficulty and an autonomous evaluation agent to dynamically select the most informative test questions based on the model's real-time performance. Specifically, AutoJudger incorporates two pivotal components: a semantic-aware retrieval mechanism to ensure that selected questions cover diverse and challenging scenarios across both vision and language modalities, and a dynamic memory that maintains contextual statistics of previously evaluated questions to guide coherent and globally informed question selection throughout the evaluation process. Extensive experiments on four representative multimodal benchmarks demonstrate that our adaptive framework dramatically reduces evaluation expenses, i.e. AutoJudger uses only 4% of the data to achieve over 90% ranking accuracy with the full benchmark evaluation on MMT-Bench.

多模态模型智能评测自适应效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。