用心理测量学方法优化大模型排行榜,让排名更科学可靠。
Improving LLM Leaderboards with Psychometrical Methodology
- 引入心理测量学方法,改进大模型性能评估的统计逻辑。
- 对比传统平均分排行,新方法能更准确反映模型真实水平。
- 适合关注模型评估公平性与可信度的研究者和开发者。
大语言模型的快速发展催生了众多评测基准,这些基准类似人类测试与问卷,旨在衡量模型在认知行为中涌现出的特性。然而,与社会科学中明确界定的能力不同,这些基准所衡量的属性往往模糊且定义不严谨。主流基准常被归入排行榜,通过简单平均得分进行模型比较,但该方法存在缺陷。本文以Hugging Face排行榜为例,对比传统平均法与心理测量学方法的排名结果。研究表明,采用心理测量技术可显著提升排行榜的稳健性与意义,为大模型评估提供更可靠的依据。
原文摘要 · Abstract (English)
The rapid development of large language models (LLMs) has necessitated the creation of benchmarks to evaluate their performance. These benchmarks resemble human tests and surveys, as they consist of sets of questions designed to measure emergent properties in the cognitive behavior of these systems. However, unlike the well-defined traits and abilities studied in social sciences, the properties measured by these benchmarks are often vaguer and less rigorously defined. The most prominent benchmarks are often grouped into leaderboards for convenience, aggregating performance metrics and enabling comparisons between models. Unfortunately, these leaderboards typically rely on simplistic aggregation methods, such as taking the average score across benchmarks. In this paper, we demonstrate the advantages of applying contemporary psychometric methodologies - originally developed for human tests and surveys - to improve the ranking of large language models on leaderboards. Using data from the Hugging Face Leaderboard as an example, we compare the results of the conventional naive ranking approach with a psychometrically informed ranking. The findings highlight the benefits of adopting psychometric techniques for more robust and meaningful evaluation of LLM performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。