arXiv:2603.19987cs.LGcs.AI2026-03被引 1

用马尔可夫状态突破大模型后训练的能力瓶颈

Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States

  • 引入显式马尔可夫状态替代历史动作序列作为状态表示
  • 在逻辑谜题任务中性能显著超越传统强化学习方法
  • 适合追求生成模型本质推理能力突破的研究者

强化学习(RL)已成为大语言模型(LLM)后训练与对齐的标准范式,但近期研究表明其存在持续的“能力天花板”:与经典RL系统能发现新策略不同,当前LLM的RL仅对预训练权重中已存在的模式进行微调。本文识别出根本性结构瓶颈:经典RL依赖紧凑、信息丰富的马尔可夫状态,而现有LLM后训练方法仍受限于不断扩展的动作历史。我们重新引入长期被忽视的经典原则——显式马尔可夫状态。理论上,我们提供了严格保证,证明利用估计的马尔可夫状态可显著降低样本复杂度。实证上,我们展示在一系列复杂逻辑谜题中引入马尔可夫状态能持续突破标准RL后训练的性能边界。结果表明,摒弃‘以历史为状态’的建模方式,转向结构化的马尔可夫表示,是解锁生成式AI开放探索与真正新推理能力的关键。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become a standard paradigm for post-training and aligning Large Language Models (LLMs), yet recent evidence suggests it faces a persistent "capability ceiling": unlike classical RL systems that discover novel strategies, RL for LLMs often acts as a mere refiner of patterns already latent in pre-trained weights. In this work, we identify a fundamental structural bottleneck: while classical RL relies on compact, informative Markov states, current LLM post-training formulations are tethered to an ever-expanding history of actions. We revisit a classical principle long central to RL yet absent from LLM post-training: explicit Markov states. Theoretically, we provide rigorous guarantees demonstrating that leveraging estimated Markov states can significantly reduce sample complexity. Empirically, we show that introducing Markov states consistently breaks the performance boundaries of standard RL post-training across a suite of complex logic puzzles. Our findings suggest that moving beyond "history-as-state" modeling in favor of structured Markovian representations is essential for unlocking open-ended discovery and genuinely new reasoning capabilities in Generative AI.

强化学习大模型训练推理能力马尔可夫状态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。