arXiv:2505.05602cs.AIstat.AP2025-05被引 16

用分层贝叶斯模型更准评估AI性能,尤其小数据下仍可靠。

HiBayES: A Hierarchical Bayesian Modeling Framework for AI Evaluation Statistics

  • 基于分层贝叶斯框架,融合广义线性模型与贝叶斯推断
  • 在每项评估少于20个样本时仍能稳定估计性能并量化不确定性
  • 适合高成本、复杂结构的AI评测,如智能体任务评估

随着大语言模型等AI系统不断发展,从固有的随机输出中稳健估计其能力,并系统量化这些估计的不确定性变得愈发重要。此外,先进的AI评估常具有嵌套的分层结构,过程复杂且测试前沿AI系统成本高昂。为此,我们提出HiBayES——一种通用的分层贝叶斯建模框架,用于AI评估统计。该框架支持在经典问答基准和高级代理式评估中进行稳健推断,尤其适用于低数据场景(如每项评估数据点少于20个)。基于广义线性模型(GLMs)、贝叶斯数据分析及正式模型比较,HiBayES提供有原则的不确定性量化和鲁棒参数估计。本文全面介绍HiBayES,包含示例、与传统统计方法的对比,以及多层级贝叶斯GLM的实用实现指导。此外,我们还提供了可直接使用的HiBayES软件包(β版)。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) and other AI systems evolve, robustly estimating their capabilities from inherently stochastic outputs while systematically quantifying uncertainty in these estimates becomes increasingly important. Further, advanced AI evaluations often have a nested hierarchical structure, exhibit high levels of complexity, and come with high costs in testing the most advanced AI systems. To address these challenges, we introduce HiBayES, a generalizable Hierarchical Bayesian modeling framework for AI Evaluation Statistics. HiBayES supports robust inferences in classical question-answer benchmarks and advanced agentic evaluations, particularly in low-data scenarios (e.g., < 20 data points per evaluation). Built on Generalized Linear Models (GLMs), Bayesian data analysis, and formal model comparison, HiBayES provides principled uncertainty quantification and robust parameter estimation. This paper offers a comprehensive introduction to HiBayES, including illustrative examples, comparisons to conventional statistical methods, and practical guidance for implementing multilevel Bayesian GLMs. Additionally, we provide a HiBayES software package [4] (Beta version) for out-of-the-box implementation.

分层贝叶斯模型评估不确定性量化小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。