大模型错误高度相关,即使架构不同也易集体犯错。
Correlated Errors in Large Language Models
- 分析350+大模型,发现错误在60%情况下同步出现
- 大型高精度模型间错误相关性依然很强
- 适合关注AI评估与算法单一风险的研究者
尽管训练数据、架构和提供商的多样性被认为可缓解大语言模型的同质化问题,但目前缺乏实证证据证明不同模型是否存在实质性差异。本文对超过350个大模型进行了大规模实证评估,使用两个主流排行榜和一份简历筛选任务。结果发现,模型错误存在显著相关性:在一个排行榜数据集上,当两个模型均出错时,它们有60%的概率同时犯相同错误。我们识别出驱动错误相关性的因素,包括共享架构和同一提供商。关键的是,即使模型具有不同架构和提供商,更大更准确的模型之间仍表现出高度错误相关性。最后,我们在两个下游任务中展示了相关性的影响:大模型作为评判者评估和招聘场景,后者印证了关于算法单一化的理论预测。
原文摘要 · Abstract (English)
Diversity in training data, architecture, and providers is assumed to mitigate homogeneity in LLMs. However, we lack empirical evidence on whether different LLMs differ meaningfully. We conduct a large-scale empirical evaluation on over 350 LLMs overall, using two popular leaderboards and a resume-screening task. We find substantial correlation in model errors -- on one leaderboard dataset, models agree 60% of the time when both models err. We identify factors driving model correlation, including shared architectures and providers. Crucially, however, larger and more accurate models have highly correlated errors, even with distinct architectures and providers. Finally, we show the effects of correlation in two downstream tasks: LLM-as-judge evaluation and hiring -- the latter reflecting theoretical predictions regarding algorithmic monoculture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。