arXiv:2504.12098cs.CL2025-04

通过区间预测评估大模型对数值问题的过度精确,发现其自信程度与实际表现严重脱节。

Gauging Overprecision in LLMs: An Empirical Study

  • 设计三阶段框架,用带置信度的区间回答测试模型自我评估能力
  • 模型在数值任务中高度不校准,置信度与区间长度无关
  • 提示技巧和任务类型影响模型精度,但答案优化效果有限

近期,大语言模型(LLMs)的过度自信问题因关乎生成内容可信度而受到广泛关注。然而,现有方法要求黑箱模型输出其置信度(即口头化置信),易受偏见和幻觉影响。受认知科学中‘过度精确’概念启发,我们设计了一个用于研究黑箱LLM过度精确性的框架,包含三个阶段:生成、精炼和评估。在生成阶段,我们以特定置信度要求模型对数值问题给出区间答案,该置信度由提示施加,无需模型自行生成。采用多种提示技术并重复多次,以考察生成过程中的随机性影响。精炼阶段对生成结果进行优化。评估阶段分析模型行为。研究发现:1)模型在数值任务上高度不可校准;2)区间长度与设定置信度无相关性,可能反映对置信概念理解不足或无法按指令调整自信;3)模型数值精确度随任务类型、答案尺度和提示方式而异;4)多数情况下,答案精炼无法提升精度。本研究为理解模型过度自信提供了新视角,并建立了过度精确性的基准。

原文摘要 · Abstract (English)

Recently, overconfidence in large language models (LLMs) has garnered considerable attention due to its fundamental importance in quantifying the trustworthiness of LLM generation. However, existing approaches prompt the \textit{black box LLMs} to produce their confidence (\textit{verbalized confidence}), which can be subject to many biases and hallucinations. Inspired by a different aspect of overconfidence in cognitive science called \textit{overprecision}, we designed a framework for its study in black box LLMs. This framework contains three main phases: 1) generation, 2) refinement and 3) evaluation. In the generation phase we prompt the LLM to generate answers to numerical questions in the form of intervals with a certain level of confidence. This confidence level is imposed in the prompt and not required for the LLM to generate as in previous approaches. We use various prompting techniques and use the same prompt multiple times to gauge the effects of randomness in the generation process. In the refinement phase, answers from the previous phase are refined to generate better answers. The LLM answers are evaluated and studied in the evaluation phase to understand its internal workings. This study allowed us to gain various insights into LLM overprecision: 1) LLMs are highly uncalibrated for numerical tasks 2) there is no correlation between the length of the interval and the imposed confidence level, which can be symptomatic of a a) lack of understanding of the concept of confidence or b) inability to adjust self-confidence by following instructions, {3) LLM numerical precision differs depending on the task, scale of answer and prompting technique 4) Refinement of answers doesn't improve precision in most cases. We believe this study offers new perspectives on LLM overconfidence and serves as a strong baseline for overprecision in LLMs.

大模型过自信数值推理可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。