arXiv:2506.17124cs.AI2025-06NeurIPS被引 6

模型无关的强化学习何时能催生推理行为?

When Can Model-Free Reinforcement Learning be Enough for Thinking?

  • 提出思想马尔可夫决策过程,形式化思考动作与策略改进的关系
  • 证明策略初始化是决定是否产生思考的关键因素
  • 验证开源大模型满足理论条件,且可推广至非语言任务

近期研究表明,大语言模型可通过无模型强化学习训练出类推理能力。这种'思考'行为既不直接带来奖励,也不改变外部状态以提高获益概率,因而其出现机制值得深入探究。本文构建了一个名为思想马尔可夫决策过程(thought MDP)的理论模型,该模型在经典MDP基础上引入抽象的思想状态与思想动作。通过该模型,证明了策略初始化对思考行为能否涌现具有决定性作用,并形式化地表明思想动作等价于代理在行动前执行一步策略改进。研究进一步发现,开源大语言模型满足理论预测的必要条件,能够实现类似思考的行为。最后,论文提出在非语言生成任务中学习思考的充分条件,并设计一个玩具场景,证明结合多任务预训练与特定思想动作可使强化学习更高效,优于无思考代理。

原文摘要 · Abstract (English)

Recent work on large language models has demonstrated the use of model-free reinforcement learning (RL) to train reasoning-like capabilities. The emergence of "thinking" through model-free RL is interesting as thinking actions neither produce reward nor change the external world state to one where the agent is more likely to get reward. This paper seeks to build a domain-independent understanding of when model-free RL will lead to such "thinking" as a strategy for reward maximization. To build this understanding, we first introduce a theoretical model which we call a thought Markov decision process (MDP). Thought MDPs minimally extend the classical MDP model to include an abstract notion of thought state and thought action. Using the thought MDP model, we prove the importance of policy initialization in determining whether or not thinking emerges and show formally that thought actions are equivalent to the agent choosing to perform a step of policy improvement before continuing to act. We then show that open-source LLMs satisfy the conditions that our theory predicts are necessary for model-free RL to produce thinking-like behavior. Finally, we hypothesize sufficient conditions that would enable thinking to be learned outside of language generation and introduce a toy domain where a combination of multi-task pre-training and designated thought actions enable more data-efficient RL compared to non-thinking agents.

强化学习大模型思考机制策略改进

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。