arXiv:2602.12424cs.CLcs.AI2026-02中稿 · ICLR

通过量化题目难度,更精准评估大模型真实能力。

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

  • 基于模型与题目双向评分机制,动态计算难度与能力。
  • 在35,550个问题上达90%人类判断一致率,优于IRT等基线。
  • 适合需要公平、高效评估模型性能的研究者使用。

基准测试为系统评估大语言模型(LLMs)性能提供了标准化框架,推动了领域进步。然而现有基准未能区分题目难度,限制了对模型能力的精细区分。为此,我们提出RankLLM,一个可量化题目难度与模型能力的新框架。该框架以难度为核心评价标准,实现模型与题目间的双向评分传播:模型答对题目则获得能力分,题目挑战能力强的模型则提升自身难度分。我们在跨多个领域的35,550个问题上评估了30个模型,结果表明RankLLM在90%程度上与人类判断一致,且持续优于IRT等强基线方法。同时具备良好稳定性、快速收敛性与高计算效率,是大规模、难度感知型大模型评估的实用方案。

原文摘要 · Abstract (English)

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field. However, existing benchmarks fail to differentiate question difficulty, limiting their ability to effectively distinguish models' capabilities. To address this limitation, we propose RankLLM, a novel framework designed to quantify both question difficulty and model competency. RankLLM introduces difficulty as the primary criterion for differentiation, enabling a more fine-grained evaluation of LLM capabilities. RankLLM's core mechanism facilitates bidirectional score propagation between models and questions. The core intuition of RankLLM is that a model earns a competency score when it correctly answers a question, while a question's difficulty score increases when it challenges a model. Using this framework, we evaluate 30 models on 35,550 questions across multiple domains. RankLLM achieves 90% agreement with human judgments and consistently outperforms strong baselines such as IRT. It also exhibits strong stability, fast convergence, and high computational efficiency, making it a practical solution for large-scale, difficulty-aware LLM evaluation.

大模型评估题目难度排序学习能力量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。