arXiv:2511.07368cs.LGcs.AI2025-11被引 3

揭示后训练中推理路径偏见的根源,解释为何探索能提升性能却仍会遗忘关键推理。

Distributional Biases in Post-Training: A Markovian Analysis of Reasoning Trajectories

  • 将推理过程建模为马尔可夫转移,区分易难步骤的概率差异
  • 证明后训练会强化高频路径,导致罕见但重要的推理被遗忘
  • 提出拒绝简单样本和KL正则化可有效保留稀有推理路径

基础模型虽具备广泛知识,但任务特定推理能力有限,促使采用基于可验证奖励的强化学习(RLVR)和测试时缩放(TTS)等后训练策略。尽管近期研究强调探索对提升pass@K的重要性,但实证发现RLVR与ORM/PRM通常强化已有路径而非拓展推理范围,引发矛盾:若未出现新模式,探索为何有效?本文借鉴Kim等(2025)视角,将简单(如化简分数)与困难(如发现对称性)推理步骤视为低概率与高概率的马尔可夫转移。在此可处理模型中,预训练对应树图发现,后训练对应思维链(CoT)重加权。理论证明RLVR与ORM/PRM均倾向于少数高概率路径,从而遗忘稀有但关键的CoT。进一步证明,如拒绝简单实例和KL正则化等探索策略有助于保留稀有CoT。实验模拟验证了理论结果。

原文摘要 · Abstract (English)

Foundation models exhibit broad knowledge but limited task-specific reasoning, motivating post-training strategies such as RL with verifiable rewards (RLVR) and test-time scaling (TTS). While recent work highlights the role of exploration in improving pass@K, empirical evidence points to a paradox: RLVR and ORM/PRM typically reinforce existing paths rather than expanding the reasoning scope, raising the question of why exploration helps if no new patterns emerge. To reconcile this paradox, we adopt the perspective of Kim et al. (2025), viewing easy (e.g., simplifying a fraction) versus hard (e.g., discovering the some symmetry) reasoning steps as low versus high probability Markov transitions. In this tractable model, pretraining corresponds to tree-graph discovering, while post-training corresponds to CoT reweighting. We provably show that, both RLVR and ORM/PRM would favor heavily to several high-probability paths, and thereby forget rare-but-crucial CoTs. Building on this, we further prove that exploration strategies such as rejecting easy instances and KL regularization help preserve rare CoTs. Empirical simulations corroborate our theoretical results.

推理机制后训练马尔可夫模型探索策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。