arXiv:2503.10079cs.CL2025-03ICCV被引 8

提出信息密度原则,评估多模态大模型评测集的实用价值

Information Density Principle for MLLM Benchmarks

  • 从谬误、难度、冗余、多样性四维度衡量评测集信息密度
  • 分析超1万样本发现新评测集提供更多信息,仍有提升空间
  • 为开发者选择与设计评测集提供可量化的判断标准

随着多模态大语言模型(MLLMs)的发展,数百个评测集被用于确保其在下游任务中的可靠性。然而,评测机制本身可能不可靠。对MLLM开发者而言,仍存在应选用哪个评测集以及测试结果是否满足需求的问题。为此,我们提出信息密度原则,评估评测集对MLLM开发所能提供的洞察力。该原则从四个关键维度进行刻画:(1) 谬误,(2) 难度,(3) 冗余,(4) 多样性。通过对超过10,000个样本的综合分析,我们量化了19个MLLM评测集的信息密度。实验表明,使用最新评测集进行测试相比旧版本能提供更多信息,但其信息密度仍有提升空间。我们希望这一原则能推动未来MLLM评测集的改进与应用。

原文摘要 · Abstract (English)

With the emergence of Multimodal Large Language Models (MLLMs), hundreds of benchmarks have been developed to ensure the reliability of MLLMs in downstream tasks. However, the evaluation mechanism itself may not be reliable. For developers of MLLMs, questions remain about which benchmark to use and whether the test results meet their requirements. Therefore, we propose a critical principle of Information Density, which examines how much insight a benchmark can provide for the development of MLLMs. We characterize it from four key dimensions: (1) Fallacy, (2) Difficulty, (3) Redundancy, (4) Diversity. Through a comprehensive analysis of more than 10,000 samples, we measured the information density of 19 MLLM benchmarks. Experiments show that using the latest benchmarks in testing can provide more insight compared to previous ones, but there is still room for improvement in their information density. We hope this principle can promote the development and application of future MLLM benchmarks. Project page: https://github.com/lcysyzxdxc/bench4bench

多模态模型评测基准信息密度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。