量化大模型评估中的三类噪声,提升实验可靠性。
Measuring all the noises of LLM Evals
- 提出全对比较法,系统测量预测、数据及总噪声
- 预测噪声通常高于数据噪声,平均可显著提效
- 帮助研究者用正确方法评估模型,避免误判
分离信号与噪声是实验的核心。将成熟统计方法有效应用于大模型评估,需关注其特有的噪声特征。本文明确界定并测量了三类噪声:给定问题下生成不同答案的预测噪声、采样问题带来的数据噪声,以及根据全方差定律得出的总噪声。为强调相对比较并提升统计功效,提出全对配对法,基于百万级题级预测,在多个评估场景中测量所有噪声成分,揭示清晰模式:第一,每次评估在所有模型对间呈现固定且高度可预测的总噪声水平;第二,配对预测噪声通常大于配对数据噪声,说明通过平均可显著降低预测噪声,大幅提升统计功效。综合测量各类噪声,使评估结果可置于上下文中评估,降低使用最优分析方法的门槛,支持更可靠的实证决策。
原文摘要 · Abstract (English)
Separating signal from noise is central to experiments. Applying well-established statistical methods effectively to LLM evals requires consideration of their unique noise characteristics. We clearly define and measure three types of noise: prediction noise from generating different answers on a given question, data noise from sampling questions, and their combined total noise following the law of total variance. To emphasize relative comparisons and gain statistical power, we propose the all-pairs paired method, which applies the paired analysis to all pairs of LLMs and measures all the noise components based on millions of question-level predictions across many evals and settings, revealing clear patterns. First, each eval exhibits a characteristic and highly predictable total noise level across all model pairs. Second, paired prediction noise typically exceeds paired data noise, which means reducing prediction noise by averaging can significantly increase statistical power. By measuring all the noises together, we can assess eval results in context, lowering the barrier of using the best analysis to make sound empirical decisions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。