解释大模型为何在训练中偷懒,导致推理能力脆弱。
The Paradox of Outcome Optimization: A Causal Information-Theoretic Bound on Reasoning Shortcuts in LLMs

- 用因果模型和信息瓶颈理论分析模型如何走捷径
- 发现数据越多,若分布单一,反而越难纠正推理错误
- 过程奖励机制能阻断简单捷径,适合需要可靠推理的场景
通过基于结果的强化学习对大型语言模型(LLM)进行对齐时,常出现一种关键失效模式:在分布内基准上表现优异,但在分布外(OOD)任务上推理能力极差。我们称此现象为‘奖励诱导流形坍缩’。本文建立一个融合结构因果模型(SCM)与信息瓶颈(IB)原理的理论框架,将推理定义为高复杂度的因果过程,而捷径学习则是利用低复杂度的虚假相关性。在随机梯度下降(SGD)的隐式归纳偏置下,只要训练分布允许对真实因果机制进行‘马尔可夫筛选’,模型优化目标就会偏向捷径解。我们提出一个新的泛化界,以语义覆盖度(η)而非样本量作为衡量标准,说明在同质分布上单纯增加数据无法修复推理缺陷。此外,过程奖励模型(PRMs)可视为拓扑过滤器,通过施加逐步互信息约束,使低复杂度的捷径流形不可行。这些结果为过程监督超越简单信用分配提供了数学基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) aligned via outcome-based Reinforcement Learning (RL) frequently exhibit a critical failure mode: they achieve high performance on in-distribution benchmarks while demonstrating brittle reasoning capabilities on out-of-distribution (OOD) tasks. We term this phenomenon Reward-Induced Manifold Collapse. We establish a theoretical framework bridging Structural Causal Models (SCM) and the Information Bottleneck (IB) principle to explain this paradox. We define reasoning as a high-complexity causal process and shortcut learning as the exploitation of low-complexity spurious correlations. Under the implicit inductive bias of Stochastic Gradient Descent (SGD), models optimized for outcome rewards are biased toward shortcut solutions whenever the training distribution allows for a ``Markovian Screening'' of the true causal mechanism. We derive a new generalization bound based on Semantic Coverage Measure ($η$) rather than sample size, showing why data scaling on homogeneous distributions may fail to correct reasoning flaws. We also show that Process Reward Models (PRMs) function as Topological Filters, enforcing step-wise mutual information constraints that render the low-complexity shortcut manifold inadmissible. These results provide a mathematical grounding for the role of process supervision beyond simple credit assignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。