arXiv:2504.02902cs.CLcs.AI2025-04被引 4

研究大模型自我改进时的置信度问题,发现越改越自信但可能更不靠谱

Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models

  • 对比三种自我改进方法,发现迭代修正导致模型越来越自信
  • 自改进后模型高置信度预测准确率下降,校准误差(ECE)持续上升
  • 在每轮改进中同步校准最有效,提升可靠性适合安全关键场景

大型语言模型(LLMs)展现出强大的自我改进能力,通过自生成反馈不断优化输出。然而,这种反思机制可能导致自偏见——即模型倾向于偏好自身先前输出。本文进一步探究其对置信度估计的影响。评估了三种典型自改进范式:基础提示、思维链(CoT)提示和基于调优的方法,发现迭代自改进会导致系统性过度自信,表现为预期校准误差(ECE)持续上升,且高置信度预测准确性下降。随后探索将置信度校准技术融入自改进过程,比较三种策略:(1)多轮自改进后校准,(2)自改进前校准,(3)每轮自改进中迭代校准。结果表明,迭代校准最有效,显著降低ECE,改善校准性能。本工作首次从校准视角系统研究自改进大模型,为平衡性能与可靠性提供重要洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable self-improvement capabilities, whereby models iteratively revise their outputs through self-generated feedback. While this reflective mechanism has shown promise in enhancing task performance, recent studies suggest that it may also introduce undesirable biases-most notably, self-bias, or the tendency of LLMs to favor their own prior outputs. In this work, we extend this line of inquiry by investigating the impact on confidence estimation. We evaluate three representative self-improvement paradigms-basic prompting, Chain-of-Thought (CoT) prompting, and tuning-based methods and find that iterative self-improvement can lead to systematic overconfidence, as evidenced by a steadily increasing Expected Calibration Error (ECE) and lower accuracy with high confidence. We then further explore the integration of confidence calibration techniques with self-improvement. Specifically, we compare three strategies: (1) applying calibration after multiple rounds of self-improvement, (2) calibrating before self-improvement, and (3) applying calibration iteratively at each self-improvement step. Our results show that iterative calibration is most effective in reducing ECE, yielding improved calibration. Our work pioneers the study of self-improving LLMs from a calibration perspective, offering valuable insights into balancing model performance and reliability.

大模型自改进置信度校准可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。