用新框架仅需3%数据,就能高效准确评估大模型能力。
Toward a unified framework for data-efficient evaluation of large language models
- 融合项目反应理论,同时支持对错和连续评分两种评估方式。
- 在5个基准上测试,用3%数据即得稳定能力估计,误差降低10%。
- 能捕捉不同任务间的结构关联,适合研究大模型评估的学者。
评估大语言模型(LLMs)在综合性基准上的表现是其发展的重要基石,但通常计算和成本高昂。尽管项目反应理论(IRT)为数据高效的评估提供了前景,现有方法受限于仅支持二元正确性指标,无法原生处理生成任务中的连续评分,并且只针对单一基准,忽略了不同指标或基准间的相关性等结构知识。为此,我们提出 LEGO-IRT,一个统一且灵活的数据高效评估框架。LEGO-IRT 的创新设计原生支持二元与连续评估指标,并引入因子化架构,显式建模并利用结构知识,将模型能力估计分解为通用成分和特定结构成分(如每项指标或每个基准)。在涵盖70个LLM和5个基准的广泛实验中,我们发现 LEGO-IRT 仅需总评估样本的3%即可获得稳定的性能估计。我们证明引入结构知识可使估计误差降低最高达10%,并揭示该框架估计的潜在能力可能更符合人类偏好。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) on comprehensive benchmarks is a cornerstone of their development, yet it's often computationally and financially prohibitive. While Item Response Theory (IRT) offers a promising path toward data-efficient evaluation by disentangling model capability from item difficulty, existing IRT-based methods are hampered by significant limitations. They are typically restricted to binary correctness metrics, failing to natively handle the continuous scores used in generative tasks, and they operate on single benchmarks, ignoring valuable structural knowledge like correlations across different metrics or benchmarks. To overcome these challenges, we introduce LEGO-IRT, a unified and flexible framework for data-efficient LLM evaluation. LEGO-IRT's novel design natively supports both binary and continuous evaluation metrics. Moreover, it introduces a factorized architecture to explicitly model and leverage structural knowledge, decomposing model ability estimates into a general component and structure-specific (e.g., per-metric or per-benchmark) components. Through extensive experiments involving $70$ LLMs across $5$ benchmarks, we show that LEGO-IRT achieves stable capability estimates using just $3\%$ of the total evaluation items. We demonstrate that incorporating structural knowledge reduces estimation error by up to $10\%$ and reveal that the latent abilities estimated by our framework may align more closely with human preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。