arXiv:2502.11028cs.CLcs.AI2025-02被引 51

通过添加干扰项提升大模型回答的可信度,减少盲目自信。

Mind the Confidence Gap: Overconfidence, Calibration, and Distractor Effects in Large Language Models

  • 用带干扰项的结构化提示,改善模型置信度与正确率的一致性。
  • 在多个数据集上,准确率最高提升460%,误差降低90%。
  • 小模型更受益,但需针对性微调和提示设计才能可靠使用。

大型语言模型在自然语言任务中表现优异,但其频繁出现的过度自信——预测置信度与真实正确性不匹配——在关键决策应用中带来显著风险。本文对九个大模型在三个事实型问答数据集上的校准问题进行了全面分析,系统比较了标准自由生成与结构化干扰项增强提示的差异。结果表明,显式引入干扰项可显著缓解校准偏差,在相对准确率上提升达460%,期望校准误差(ECE)下降最多90%。尽管整体趋势明显,仍发现复杂现象:大规模经过强化学习人类反馈(RLHF)微调的模型虽具内在校准优势,但在简单问题上反而出现更严重的校准偏差;而较小模型虽从干扰提示中获益更大,但仍显著存在校准问题。通过对不同题型的深入分析,发现人物类问题存在持续性的校准失败。最后提出具体建议:针对性微调、结构化提示设计和策略性模型选择,以确保大模型部署的可靠性与可信性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show remarkable proficiency in natural language tasks, yet their frequent overconfidence-misalignment between predicted confidence and true correctness-poses significant risks in critical decision-making applications. We present a comprehensive analysis on calibration in LLMs across nine LLMs and three factual Question-Answering (QA) datasets, systematically comparing standard free-generation settings against structured distractor-augmented prompts. Our evaluation reveals that explicitly incorporating distractors can substantially mitigate miscalibration, achieving relative accuracy improvements up to 460% and ECE reductions up to 90%. Despite general trends, we uncover nuanced findings: large RLHF-tuned models display inherent calibration strengths but can paradoxically suffer increased miscalibration on easier queries, whereas smaller models benefit disproportionately from distractor prompts but remain significantly miscalibrated. Through detailed analyses across question types, we identify persistent calibration failures, particularly in person-based queries. We conclude with concrete recommendations-targeted fine-tuning, structured prompting, and strategic model choice-to ensure reliable, trustworthy LLM deployments.

大模型校准提示工程可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。