arXiv:2505.13774cs.AI2025-05NeurIPS被引 20

提出新方法评估大模型推理过程是否可信

Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models

  • 通过反事实干预测试推理步骤间的因果关系
  • 发现当前模型推理过程常与最终答案不一致
  • 适合关注模型可解释性与可靠性研究者

大型推理模型(LRMs)通过引入思维草稿,实现多路径链式思考以解决复杂问题。确保中间推理过程的忠实性对可靠监控、解释和有效控制至关重要。本文提出系统性的反事实干预框架,从两个维度评估思维草稿的忠实性:(1) 草稿内忠实性,通过反事实插入推理步骤,检验其是否对后续步骤和最终结论产生因果影响;(2) 草稿到答案忠实性,通过扰动草稿结尾逻辑,评估最终答案是否与其保持逻辑一致并依赖于草稿。我们在六种先进LRM上进行广泛实验,结果表明当前模型对中间推理步骤表现出选择性忠实,且频繁无法忠实对应草稿结论。该发现凸显了提升先进推理模型推理过程忠实性与可解释性的必要性。

原文摘要 · Abstract (English)

Large Reasoning Models (LRMs) have significantly enhanced their capabilities in complex problem-solving by introducing a thinking draft that enables multi-path Chain-of-Thought explorations before producing final answers. Ensuring the faithfulness of these intermediate reasoning processes is crucial for reliable monitoring, interpretation, and effective control. In this paper, we propose a systematic counterfactual intervention framework to rigorously evaluate thinking draft faithfulness. Our approach focuses on two complementary dimensions: (1) Intra-Draft Faithfulness, which assesses whether individual reasoning steps causally influence subsequent steps and the final draft conclusion through counterfactual step insertions; and (2) Draft-to-Answer Faithfulness, which evaluates whether final answers are logically consistent with and dependent on the thinking draft, by perturbing the draft's concluding logic. We conduct extensive experiments across six state-of-the-art LRMs. Our findings show that current LRMs demonstrate selective faithfulness to intermediate reasoning steps and frequently fail to faithfully align with the draft conclusions. These results underscore the need for more faithful and interpretable reasoning in advanced LRMs.

推理模型可解释性忠实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。