arXiv:2603.09985cs.CLcs.AI2026-03被引 1

研究发现大模型像人一样,水平越差越自信,存在认知偏差。

The Dunning-Kruger Effect in Large Language Models: An Empirical Study of Confidence Calibration

  • 通过2.4万次实验测试四款大模型的自信心评估能力
  • 表现差的模型(如Kimi K2)错误率超76%却极度自信
  • 部分模型校准度极差,不适合高风险场景使用

大型语言模型在多种任务中展现出卓越能力,但其自我信心评估的准确性仍不明确。我们开展了一项实证研究,考察大模型是否表现出类似达尼尔-克鲁格效应——即能力不足者往往过度高估自身水平。在总计24,000次实验的四个基准数据集上,评估了四款先进模型(Claude Haiku 4.5、Gemini 2.5 Pro、Gemini 2.5 Flash 和 Kimi K2)。结果揭示显著的校准差异:Kimi K2 的准确率仅为23.3%,却表现出严重过自信,预期校准误差(ECE)高达0.726;而 Claude Haiku 4.5 准确率达75.4%,校准最佳(ECE = 0.122)。这些发现表明,性能较差的模型表现出明显更高的过自信,这一现象与人类认知中的达尼尔-克鲁格效应高度相似。研究讨论了对大模型在高风险应用中安全部署的启示。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet their ability to accurately assess their own confidence remains poorly understood. We present an empirical study investigating whether LLMs exhibit patterns reminiscent of the Dunning-Kruger effect -- a cognitive bias where individuals with limited competence tend to overestimate their abilities. We evaluate four state-of-the-art models (Claude Haiku 4.5, Gemini 2.5 Pro, Gemini 2.5 Flash, and Kimi K2) across four benchmark datasets totaling 24,000 experimental trials. Our results reveal striking calibration differences: Kimi K2 exhibits severe overconfidence with an Expected Calibration Error (ECE) of 0.726 despite only 23.3% accuracy, while Claude Haiku 4.5 achieves the best calibration (ECE = 0.122) with 75.4% accuracy. These findings demonstrate that poorly performing models display markedly higher overconfidence -- a pattern analogous to the Dunning-Kruger effect in human cognition. We discuss implications for safe deployment of LLMs in high-stakes applications.

大模型自信评估校准误差认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。