arXiv:2504.06564cs.CL2025-04EMNLP被引 14

研究大模型如何自我判断对错,发现越会推理越容易盲目自信。

Thinking Out Loud: Do Reasoning Models Know When They're Right?

  • 通过分析模型自述的把握度,观察其自我反思能力。
  • 微调和强化学习能逐步提升自述准确性,但小模型更难承认不知道。
  • 推理越简略,模型越自信,存在过度自信风险,适合关注模型可信度的研究者。

大型推理模型(LRMs)在复杂推理任务中展现出惊人能力,依赖增加的推理时间计算,并表现出类似人类自我反思的行为。尽管这些模型具备显著的自我反思潜力,其与其它行为之间的互动仍不清晰。本文以模型自述的置信度为切入点,探究其自我反思的本质。研究发现,在推理链上进行监督微调(即知识蒸馏)和强化学习,可使模型在高推理强度任务中逐步改善自述校准能力。然而,结果也显示,推理模型可能对其知识边界认知减弱,体现在事实性基准测试中“我不知道”响应率显著降低。此外,我们分析了自述置信度与推理链长度的关系,发现模型在提供更短或更简单的推理时,往往表达更高置信度。这些发现表明,以推理为导向的训练虽能提升任务表现,但可能带来‘推理代价’:小模型对自身知识局限的认知能力下降。更广泛而言,这种知识边界的模糊化会损害模型的可靠性——模型变得越来越自信,却未同步理解何时应保持沉默。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) have recently demonstrated impressive capabilities in complex reasoning tasks by leveraging increased test-time computation and exhibiting behaviors reminiscent of human-like self-reflection. While LRMs show a clear capacity for valuable self-reflection, how this ability interacts with other model behaviors remains underexplored. We investigate this connection by analyzing verbalized confidence, how models articulate their certainty, as a lens into the nature of self-reflection in LRMs. We find that supervised fine-tuning on reasoning traces (i.e., distillation) and reinforcement learning can improve verbalized calibration in reasoning-intensive settings in a progressive, laddered fashion. However, our results also indicate that reasoning models may possess a diminished awareness of their own knowledge boundaries, as evidenced by significantly lower "I don't know" response rates on factuality benchmarks. Moreover, we examine the relationship between verbalized confidence and reasoning chains, finding that models tend to express higher confidence when providing shorter or less elaborate reasoning. Our findings highlight how reasoning-oriented training can enhance performance in reasoning-centric tasks while potentially incurring a "reasoning tax," a cost reflected in the model's reduced ability to accurately recognize the limits of its own knowledge in small-scale models. More broadly, our work showcases how this erosion of knowledge boundaries can compromise model faithfulness, as models grow more confident without a commensurate understanding of when they should abstain.

大模型推理可信度自省

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。