arXiv:2609.04194cs.CLcs.LG2026-09

模型推理步骤的可读性不等于可解释性,文本本身难反映关键步骤。

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

论文配图:Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
图 1 · 摘自论文原文
  • 用蒙特卡洛模拟量化每步推理对结果的实际贡献
  • 大模型判断重要步骤的能力远低于理论上限
  • 即使微调为评论器,正确答案的步骤仍难识别

链式思维模型的推理过程看似清晰可读,被广泛用于诊断错误、评估忠实度及构建过程奖励模型。但这些方法假设每一步文字能反映其功能重要性。我们以步骤的‘优势’(即包含该步后预期奖励的提升)作为真实重要性的度量,通过蒙特卡洛滚动生成。结果显示,尽管高能力大模型优于随机基准,但仍远低于噪声上限;将模型微调为步骤级评论器虽对错误回答有显著提升,但在正确回答中仍无法接近理论上限,表明推理步骤的重要性仅部分可从文本中恢复。研究警示:不应将推理过程的可读性等同于可解释性,尤其对过程奖励建模具有重要意义。

原文摘要 · Abstract (English)

Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.

可解释性链式思维大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。