arXiv:2506.03723cs.CLcs.AI2025-06被引 7

让大模型自己检查答案,不用额外训练就能更准更可信。

Verbalized Confidence Triggers Self-Verification: Emergent Behavior Without Explicit Reasoning Supervision

  • 用标量信心值微调,模型自发生成自我验证的推理过程。
  • 低信心时输出更长自检内容,高信心时更简洁,提升准确率与可解释性。
  • 无需复杂监督或强化学习,适合追求可信AI的开发者使用。

不确定性校准对大型语言模型的安全部署至关重要,尤其当用户依赖模型的口语化信心评估时。尽管已有研究聚焦于分类器或简短生成,链式思维(CoT)推理中的信心校准仍基本未被探索。令人惊讶的是,仅使用标量信心标签进行监督微调,即可激发模型产生自我验证行为,而无需任何显式的推理监督或基于强化学习的奖励机制。即使训练时未提供任何自验证示例,模型也能在低信心问题上生成更长、自我检查的回应,在高信心问题上则给出更简洁的答案。我们进一步提出一种简单的测试时缩放重思考方法,通过校准后的不确定性提升性能。在GSM8K及独立推理任务如MATH-500和ARC-Challenge上的实验表明,该信心感知微调方法不仅提升了校准度与准确性,还通过使推理路径与信心水平对齐增强了可解释性。

原文摘要 · Abstract (English)

Uncertainty calibration is essential for the safe deployment of large language models (LLMs), particularly when users rely on verbalized confidence estimates. While prior work has focused on classifiers or short-form generation, confidence calibration for chain-of-thought (CoT) reasoning remains largely unexplored. Surprisingly, we find that supervised fine-tuning with scalar confidence labels alone suffices to elicit self-verification behavior of language models, without any explicit reasoning supervision or reinforcement learning-based rewards. Despite being trained only to produce a verbalized confidence score without any self-verifying examples, the model learns to generate longer and self-checking responses for low-confidence queries while providing more concise answers for high-confidence ones. We further propose a simple rethinking method that boosts performance via test-time scaling based on calibrated uncertainty. Experiments on GSM8K and held-out reasoning tasks such as MATH-500 and ARC-Challenge show that our confidence-aware fine-tuning improves both calibration and accuracy, while also enhancing interpretability by aligning the model's reasoning path with its confidence.

大模型可信推理自验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。