arXiv:2510.05457cs.AIcs.CL2025-10被引 2

发现代码大模型也存在能力越低越自信的自我错觉。

Do Code Models Suffer from the Dunning-Kruger Effect?

  • 用模型自信心与实际表现对比,检测其认知偏差
  • 低能力模型和冷门语言任务中过自信现象更明显
  • 提示开发者警惕模型在陌生场景下的误判

随着人工智能系统在创意与技术领域越来越多地与人类协作,关于认知边界与偏见的问题日益突出。本文研究了大语言模型在编程任务中是否存在达克效应(Dunning-Kruger Effect, DKE)——即能力有限者高估自身水平的现象。通过分析模型在多种编程语言中的自信度与真实性能,我们发现AI模型表现出与人类相似的过度自信模式,尤其在不熟悉或资源稀缺的编程领域更为显著。实验表明,能力较低的模型以及在罕见编程语言上运行的模型表现出更强的类似达克效应的偏差,说明该偏差强度与模型能力呈正相关。

原文摘要 · Abstract (English)

As artificial intelligence systems increasingly collaborate with humans in creative and technical domains, questions arise about the cognitive boundaries and biases that shape our shared agency. This paper investigates the Dunning-Kruger Effect (DKE), the tendency for those with limited competence to overestimate their abilities in state-of-the-art LLMs in coding tasks. By analyzing model confidence and performance across a diverse set of programming languages, we reveal that AI models mirror human patterns of overconfidence, especially in unfamiliar or low-resource domains. Our experiments demonstrate that less competent models and those operating in rare programming languages exhibit stronger DKE-like bias, suggesting that the strength of the bias is proportionate to the competence of the models.

大模型认知偏差代码生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。