arXiv:2604.23333cs.LGcs.CL2026-04被引 2

让大模型推理更可信,避免过度自信导致的幻觉

Process Supervision of Confidence Margin for Calibrated LLM Reasoning

论文配图:Process Supervision of Confidence Margin for Calibrated LLM Reasoning
图 1 · 摘自论文原文
  • 用信心差距作为中间步骤的奖励信号,提升推理过程的可靠性
  • 在数学、代码等任务上显著改善模型校准度,准确率不降反升
  • 适合需要可靠置信度的场景,如安全关键应用或模型集成

通过强化学习扩展测试时计算已成为提升大语言模型推理能力的可靠路径。然而,基于结果的奖励常导致模型过度自信,引发幻觉、不可靠的置信度控制以及不必要的计算资源分配。我们提出信心差距强化学习(RLCM),一种感知校准的强化学习框架,通过在中间预算生成中采用增强型过程奖励,联合优化正确性与置信度可靠性。不同于将置信度对齐于正确概率,RLCM鼓励模型在单条推理轨迹中扩大正确与错误步骤之间的信心差距。在数学、代码、逻辑和科学基准测试中,该方法显著提升了校准性能,同时保持或提高了准确率。进一步表明,具备校准置信度信号的模型可实现更高效的同调风险控制和有效的置信度加权聚合。

原文摘要 · Abstract (English)

Scaling test-time computation with reinforcement learning (RL) has emerged as a reliable path to improve large language models (LLM) reasoning ability. Yet, outcome-based reward often incentivizes models to be overconfident, leading to hallucinations, unreliable confidence-based control, and unnecessary compute allocation. We introduce Reinforcement Learning with Confidence Margin (RLCM), a calibration-aware RL framework that jointly optimizes correctness and confidence reliability via a margin-enhanced process reward over intermediate-budget completions. Rather than aligning confidence to correctness likelihoods, RLCM encourages the model to widen the confidence margin between correct and incorrect steps within a single reasoning trajectory. Across mathematical, code, logic and science benchmarks, our method substantially improves calibration while maintaining or improving accuracy. We further show that, with calibrated confidence signals, the resulting models enable more efficient conformal risk control and effective confidence-weighted aggregation

大模型推理置信度校准强化学习可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。