arXiv:2603.22816cs.CLcs.AI2026-03

提出新指标与训练法,让大模型真正依赖推理步骤作答

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness

  • 用SLRC衡量推理步骤是否被真实使用,证明其具因果一致性
  • 发现高可靠模型更易讨好用户,提出整合忠诚度的RIS评分
  • 新训练法LC-CoSR显著降低推理僵化,无需外部模型依赖

语言模型越来越通过分步推理来展示思考过程,但这些步骤是否真正被使用,还是答案早已固定?我们提出步级推理能力(SLRC)度量,并证明其为一致的因果估计器(定理1)。我们提出具有李雅普诺夫稳定性保障的LC-CoSR训练方法,直接缓解推理僵化。在16个前沿模型(o4-mini、GPT-5.4、Claude Opus、Grok-4、DeepSeek-R1、Gemini 2.5 Pro等)上,覆盖六个领域,样本量N=133-500,发现推理可分为三种模式。OpenAI的o4-mini在五个任务中表现出74-88%的步骤必要性(73.8-88.3%),是本研究中最高SLRC值。关键差异在于基于强化学习的推理训练,而非思维令牌数量:Grok-4的推理模式比非推理模式更不忠实(1.4% vs 7.2%必要性)。我们发现忠诚度悖论——高SLRC模型更易讨好,提出推理完整性评分(RIS = SLRC × (1−讨好率)),显著预测错误检测能力(rho=0.66, p=0.026)。LC-CoSR相比FARL和CSR基线,负奖励减少2.6倍,且无需外部模型依赖。

原文摘要 · Abstract (English)

Language models increasingly show their work by writing step-by-step reasoning before answering. But are these steps genuinely used, or is the answer rigid - fixed before reasoning begins? We introduce the Step-Level Reasoning Capacity (SLRC) metric and prove it is a consistent causal estimator (Theorem 1). We propose LC-CoSR, a training method with Lyapunov stability guarantees that directly reduces rigidity. Evaluating 16 frontier models (o4-mini, GPT-5.4, Claude Opus, Grok-4, DeepSeek-R1, Gemini 2.5 Pro, and others) across six domains at N=133-500, we find reasoning falls into three modes. OpenAI's o4-mini shows 74-88% step necessity on five of six tasks (73.8-88.3%) - the highest SLRC in our study. The critical differentiator is RL-based reasoning training, not thinking tokens: Grok-4's reasoning mode shows lower faithfulness than its non-reasoning mode (1.4% vs 7.2% necessity). We discover a faithfulness paradox - high-SLRC models are more susceptible to sycophancy - and propose the Reasoning Integrity Score (RIS = SLRC x (1-Sycophancy)), which significantly predicts error detection (rho=0.66, p=0.026). LC-CoSR achieves 2.6x less negative reward than FARL and CSR baselines without external model dependencies.

推理评估模型可靠性强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。