LLM评估中的隐藏误差会扭曲结果,影响模型选择与结论可信度。
Hidden Measurement Error in LLM Pipelines Distorts Annotation, Evaluation, and Benchmarking
- 分解评估管道中各类不确定性来源,区分可缩小的方差与设计敏感项。
- 修正后标准误比传统方法大40%-60%,95%置信区间覆盖率稳定在目标水平。
- 小规模预实验可预测优化方向,低成本降低评估误差并提升人类一致性。
LLM评估决定了哪些模型被部署、安全标准如何制定、研究结论能否发表以及对AI劳动力影响的预测。然而,标准置信区间忽略了评判模型选择、模型温度和提示语措辞带来的变异性,导致覆盖不足,且随数据量增加而恶化。遗漏的方差足以逆转结论;未对这些因素取平均的评估流程,为“基准操纵”提供了可乘之机。本文将LLM评估中的不确定性分解为不同来源,区分随数据增多而减小的方差与受研究者设计选择影响的敏感性,并通过设计研究投影减少总评估误差(TEE)。在多个实验中,未经校正的标准误比经TEE校正的低40%-60%。基于Chatbot Arena数据,我们发现:随着样本量n增大,传统95%置信区间覆盖率下降,而TEE校正后保持在95%;校正后的流程将基准操控空间从56降至32 Elo(K=27),低于人类排行榜基线。进一步表明,小规模预实验即可恢复可信置信区间,并预测何种设计改进最能提升精度。据此调整方案,在同等成本下使MMLU估计误差减半,同时在Chatbot Arena上与人类投票的一致性提升7.9个百分点。
原文摘要 · Abstract (English)
LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made. Yet standard confidence intervals ignore variability from judge model choice, model temperature, and prompt phrasing, producing under-coverage that worsens with more data. The omitted variance can shift results enough to reverse conclusions \citep{baumann2025llmhacking, huang2026dropping}; pipelines that fail to average over it leave the surface that ``benchmark hacking'' exploits \citep{singh2025leaderboard}. This paper decomposes LLM pipeline uncertainty into its sources, distinguishes variance that shrinks with more data from sensitivity to researcher design choices, and uses design-study projections to reduce total evaluation error (TEE). Across the demonstrations, naive standard errors are 40 - 60\% smaller than the TEE-corrected SE. Using Chatbot Arena data, we show naive 95\% CI coverage drops as $n$ grows while TEE-corrected coverage holds at 95\%, and TEE-guided pipelines restrict the benchmark gaming surface from 56 to 32 Elo ($K=27$), below the human-leaderboard baseline. We show further that a small pilot recovers honest CIs and projects which design changes most improve precision. Acting on those projections halves MMLU estimation error against the answer key at equivalent cost, and raises per-match agreement with human votes by 7.9 percentage points on Chatbot Arena.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。