arXiv:2410.03492cs.CL2024-10被引 34

量化大模型评测分数的不确定性,提升评估可复现性。

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

  • 通过重复实验测量评分波动,评估模型结果的不确定性
  • 发现即使固定随机种子,模型仍存在显著评分差异
  • 提出低成本方法,适合关注评测可靠性的研究者

大型语言模型具有随机性,即便在温度设为零且固定随机种子的情况下,也未必产生确定性输出。然而,多数基准测试研究未量化这种不确定性,部分原因在于重复实验耗时且成本高。本文利用用于测试模型推理方位能力的基准,探究重复实验对平均得分与预测区间的影响。提出一种高效、低成本的不确定性量化方法,并就可复现的大模型评估提出建议。

原文摘要 · Abstract (English)

Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark studies attempt to quantify uncertainty, partly due to the time and cost of repeated experiments. We use benchmarks designed for testing LLMs' capacity to reason about cardinal directions to explore the impact of experimental repeats on mean score and prediction interval. We suggest a simple method for cost-effectively quantifying the uncertainty of a benchmark score and make recommendations concerning reproducible LLM evaluation.

大模型评估不确定性可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。