arXiv:2509.03646cs.AIcs.CL2025-09被引 43

通过强化学习让大模型自发形成分层推理能力

Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning

  • 发现模型在强化学习中会自然发展出高层策略与底层执行的分层结构
  • 提出HICRA算法,聚焦高价值规划令牌优化,显著提升推理性能
  • 适合关注大模型认知机制与高效训练方法的研究者

强化学习(RL)已被证明能有效提升大语言模型(LLMs)的复杂推理能力,但其内在机制仍不清晰。我们的分析表明,诸如‘顿悟时刻’、‘长度缩放’和熵动态等看似孤立的现象,实则均为一种涌现的推理层级结构的表现,类似于人类认知中高层战略规划与底层执行的分离。我们揭示了一种关键的双阶段动态:初期模型受限于程序正确性,需先提升底层技能;随后学习瓶颈发生转移,性能提升主要来自对高层战略规划的探索与掌握。这一发现暴露了现有RL算法(如GRPO)的固有缺陷——优化压力均匀分布于所有标记,稀释了学习信号。为此,我们提出层次感知信用分配(HICRA),将优化集中在高影响力规划标记上。大量实验验证,HICRA显著优于强基线,并为推理能力的演进提供了深刻洞察。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has proven highly effective at enhancing the complex reasoning abilities of Large Language Models (LLMs), yet underlying mechanisms driving this success remain largely opaque. Our analysis reveals that puzzling phenomena like ``aha moments", ``length-scaling'' and entropy dynamics are not disparate occurrences but hallmarks of an emergent reasoning hierarchy, akin to the separation of high-level strategic planning from low-level procedural execution in human cognition. We uncover a compelling two-phase dynamic: initially, a model is constrained by procedural correctness and must improve its low-level skills. The learning bottleneck then decisively shifts, with performance gains being driven by the exploration and mastery of high-level strategic planning. This insight exposes a core inefficiency in prevailing RL algorithms like GRPO, which apply optimization pressure agnostically and dilute the learning signal across all tokens. To address this, we propose Hierarchy-Aware Credit Assignment (HICRA), an algorithm that concentrates optimization efforts on high-impact planning tokens. Our extensive experiments validate that HICRA significantly outperforms strong baselines, and offer deep insights into how reasoning advances through the lens of strategic exploration.

大模型推理强化学习分层结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。