arXiv:2410.05222cs.LGcs.CL2024-10EMNLP被引 3

用少量数据精准评估大模型在特定主题的表现

Precise Model Benchmarking with Only a Few Observations

  • 采用经验贝叶斯方法融合直接估计与回归估计
  • 在小样本子集上显著降低均方误差,提升精度
  • 适合需要高精度评估小众主题的模型研究者

如何精确估计大语言模型(LLM)在大型问答数据集中特定主题上的准确率?标准的直接估计法(对每个子组内问题的准确率取平均)在小样本子组上方差过大;而合成回归建模虽利用其他主题信息,却可能产生偏差,导致大子组估计不可靠。本文提出一种简单有效的经验贝叶斯(EB)估计器,为每个子组分别平衡直接估计与回归估计,显著提升子组级性能估计的精度。多个数据集上的实验表明,该方法相较直接法与回归法均能持续获得更精确的估计,均方误差大幅降低。EB估计的置信区间具有接近名义覆盖率,且比直接估计更窄。在表格数据和视觉数据上的额外实验也验证了该方法的优势。

原文摘要 · Abstract (English)

How can we precisely estimate a large language model's (LLM) accuracy on questions belonging to a specific topic within a larger question-answering dataset? The standard direct estimator, which averages the model's accuracy on the questions in each subgroup, may exhibit high variance for subgroups (topics) with small sample sizes. Synthetic regression modeling, which leverages the model's accuracy on questions about other topics, may yield biased estimates that are too unreliable for large subgroups. We prescribe a simple yet effective solution: an empirical Bayes (EB) estimator that balances direct and regression estimates for each subgroup separately, improving the precision of subgroup-level estimates of model performance. Our experiments on multiple datasets show that this approach consistently provides more precise estimates of the LLM performance compared to the direct and regression approaches, achieving substantial reductions in the mean squared error. Confidence intervals for EB estimates also have near-nominal coverage and are narrower compared to those for the direct estimator. Additional experiments on tabular and vision data validate the benefits of this EB approach.

模型评估小样本经验贝叶斯大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。