测试大模型推理时的自信程度,发现它们常高估自己且改答案后更错。
Confidence in the Reasoning of Large Language Models
- 用反复提问和自评分数评估模型自信度
- 改答案后准确率反而下降,自评信心普遍偏高
- 模型缺乏一致的内在自信机制,提示词影响很大
大型语言模型在推理方面研究日益增多,但对其回答不确定性的讨论仍不足。本文旨在评估大模型对自己答案的自信程度及其与准确性的关系。自信度通过两种方式衡量:(i) 要求重新思考时是否坚持原答案(定性);(ii) 模型自报的信心分值(定量)。我们考察了 GPT4o、GPT4-turbo 与 Mistral 三种模型在因果判断、形式谬误以及概率统计谜题与悖论三个基准数据集上的表现。尽管模型性能显著优于随机猜测,但其修改初始答案的倾向差异巨大。定性自信与准确率呈正相关,但第二次回答的准确率通常低于第一次。模型普遍存在自评信心过高现象,且自信度仅部分由底层词元概率解释。提示词对定性自信有显著影响,加上过度自信倾向表明当前大模型并无内在一致的自信认知机制。
原文摘要 · Abstract (English)
There is a growing literature on reasoning by large language models (LLMs), but the discussion on the uncertainty in their responses is still lacking. Our aim is to assess the extent of confidence that LLMs have in their answers and how it correlates with accuracy. Confidence is measured (i) qualitatively in terms of persistence in keeping their answer when prompted to reconsider, and (ii) quantitatively in terms of self-reported confidence score. We investigate the performance of three LLMs -- GPT4o, GPT4-turbo and Mistral -- on two benchmark sets of questions on causal judgement and formal fallacies and a set of probability and statistical puzzles and paradoxes. Although the LLMs show significantly better performance than random guessing, there is a wide variability in their tendency to change their initial answers. There is a positive correlation between qualitative confidence and accuracy, but the overall accuracy for the second answer is often worse than for the first answer. There is a strong tendency to overstate the self-reported confidence score. Confidence is only partially explained by the underlying token-level probability. The material effects of prompting on qualitative confidence and the strong tendency for overconfidence indicate that current LLMs do not have any internally coherent sense of confidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。