大模型自信程度常高于实际表现,难题更易过度自信。
Confidence Calibration in Large Language Models

- 通过LifeEval测试评估模型在不同难度下的自信度
- 模型平均自信度高于准确率,难题下过度自信最严重
- 适合关注模型可靠性与可信AI的研究者
我们研究了大语言模型(LLMs)在多种任务中的置信度校准问题。一项预先注册的研究结果显示,当前主流的LLMs和人类一样,过于自信:整体上,其置信度高于实际准确率。值得注意的是,这种过度自信现象受‘难易效应’显著调节——在困难任务上过度自信最为明显;相反,在简单任务上反而出现明显的低估。为此,我们提出了LifeEval,一个用于评估模型在不同难度水平下校准性能的测试基准。
原文摘要 · Abstract (English)
We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the current crop of LLMs are, like people, too sure they are right: confidence exceeds accuracy, on average. Importantly, however, this tendency is moderated by a powerful hard-easy effect, wherein overconfidence is greatest on difficult tests; by contrast, easy tests actually show substantial underconfidence. We develop LifeEval, a test for evaluating model calibration across levels of difficulty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。