高分模型也可能严重分歧,影响科研结果可靠性
Benchmark Illusion: Disagreement among LLMs and Its Scientific Consequences
- 用两个推理基准测试发现:分数相近的模型在16%-66%题目上答案不同
- 替换标注模型可使教育与政治学研究效应估计变化超80%,甚至反转方向
- 提醒科研人员警惕模型选择带来的隐藏偏差,尤其在数据标注场景
基准测试是衡量大语言模型(LLM)进展和可信度的基础。然而我们的分析显示,基准准确率的表面收敛可能掩盖深层的认知分歧。基于MMLU-Pro和GPQA两个主要推理基准,我们发现即使准确率相近,不同模型在16%-66%的题目上仍存在分歧,其中顶尖模型间分歧率达16%-38%。这些差异表明各模型具有不同的错误特征。当此类模型用于科学数据标注与推断时,其隐含分歧会传递至研究结果:在对已发表教育学与政治学研究的重分析中,更换标注模型可使估算处理效应变化超过80%,部分情况下甚至改变符号。这揭示了‘基准幻象’现象——看似一致的准确率背后可能隐藏着显著分歧,模型选择成为影响科学可复现性的隐蔽但关键变量。
原文摘要 · Abstract (English)
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence. Using two major reasoning benchmarks - MMLU-Pro and GPQA - we show that LLMs achieving comparable accuracy still disagree on 16-66% of items, and 16-38% among top-performing frontier models. These discrepancies suggest distinct error profiles for different LLMs. When such models are used for scientific data annotation and inference, their hidden disagreements propagate into research results: in re-analyses of published studies in education and political science, switching the annotation model can change estimated treatment effects by more than 80%, and in some cases reverses their sign. Together, these findings illustrate a benchmark illusion, where equal accuracy may conceal disagreement, with model choice becoming a hidden yet consequential variable for scientific reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。