arXiv:2510.24505cs.CL2025-10被引 6

用自然语言批评提升大模型置信度校准,更可信。

CritiCal: Can Critique Help LLM Uncertainty or Confidence Calibration?

  • 用自然语言批评来优化置信度,比直接调数值更有效。
  • 在复杂推理任务中超越GPT-4o,且泛化能力强。
  • 适合高风险场景下需要可靠置信度的用户。

大型语言模型(LLMs)的置信度校准对高风险领域的安全应用至关重要,清晰的置信表达能增强用户信任。传统方法依赖模仿参考置信表述,但难以捕捉准确评估置信度所需的推理过程。本文提出使用自然语言批评作为解决方案,因其更适合置信度校准——精确的黄金置信标签难以获取,常需多次生成。研究探讨两个核心问题:(1)应批评什么?不确定性(以问题为中心)或置信度(以答案为核心)?分析表明,置信度适用于多选任务,而不确定性在开放问答中表现更优。(2)如何批评?自评或批评校准训练?本文提出Self-Critique,使模型能自我批评并优化置信度;以及CritiCal,一种新颖的批评校准训练方法,利用自然语言批评改进置信度校准,摆脱直接数值优化。实验显示,CritiCal显著优于Self-Critique及其他基线,甚至在复杂推理任务中超越教师模型GPT-4o。CritiCal在分布外设置下也表现出强鲁棒性,推动了大模型可靠性的发展。

原文摘要 · Abstract (English)

Accurate confidence calibration in Large Language Models (LLMs) is critical for safe use in high-stakes domains, where clear verbalized confidence enhances user trust. Traditional methods that mimic reference confidence expressions often fail to capture the reasoning needed for accurate confidence assessment. We propose natural language critiques as a solution, ideally suited for confidence calibration, as precise gold confidence labels are hard to obtain and often require multiple generations. This paper studies how natural language critiques can enhance verbalized confidence, addressing: (1) What to critique: uncertainty (question-focused) or confidence (answer-specific)? Analysis shows confidence suits multiple-choice tasks, while uncertainty excels in open-ended scenarios. (2) How to critique: self-critique or critique calibration training? We propose Self-Critique, enabling LLMs to critique and optimize their confidence beyond mere accuracy, and CritiCal, a novel Critique Calibration training method that leverages natural language critiques to improve confidence calibration, moving beyond direct numerical optimization. Experiments show that CritiCal significantly outperforms Self-Critique and other competitive baselines, even surpassing its teacher model, GPT-4o, in complex reasoning tasks. CritiCal also shows robust generalization in out-of-distribution settings, advancing LLM's reliability.

置信度校准大模型自然语言批评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。