模型无关强化学习智能体能自发形成内部规划,与人类推理相似。
Interpreting Emergent Planning in Model-Free Reinforcement Learning
- 通过概念可解释性方法挖掘智能体内部的计划表征
- 发现智能体在解谜任务中具备预测长期动作影响的能力
- 其规划机制类似并行双向搜索,适合研究大模型推理机制
我们首次提供了机制性证据,表明模型无关强化学习智能体能够学会规划。该研究基于概念可解释性方法,对Guez等人(2019)提出的通用模型无关智能体DRC在经典规划基准任务Sokoban中的行为进行分析。结果表明,该智能体利用习得的概念表征,在内部形成既能预测动作对环境的长期影响,又能指导动作选择的计划。研究方法包括:(1) 检测与规划相关的概念;(2) 探查智能体表征中的计划形成过程;(3) 通过干预验证所发现计划对行为的因果影响。此外,计划能力的出现与测试时可扩展计算带来的收益提升同步发生。定性分析显示,智能体学到的规划算法与并行双向搜索高度相似。这些发现深化了对强化学习智能体内部规划机制的理解,对当前大模型通过强化学习涌现出推理能力的研究具有重要意义。
原文摘要 · Abstract (English)
We present the first mechanistic evidence that model-free reinforcement learning agents can learn to plan. This is achieved by applying a methodology based on concept-based interpretability to a model-free agent in Sokoban -- a commonly used benchmark for studying planning. Specifically, we demonstrate that DRC, a generic model-free agent introduced by Guez et al. (2019), uses learned concept representations to internally formulate plans that both predict the long-term effects of actions on the environment and influence action selection. Our methodology involves: (1) probing for planning-relevant concepts, (2) investigating plan formation within the agent's representations, and (3) verifying that discovered plans (in the agent's representations) have a causal effect on the agent's behavior through interventions. We also show that the emergence of these plans coincides with the emergence of a planning-like property: the ability to benefit from additional test-time compute. Finally, we perform a qualitative analysis of the planning algorithm learned by the agent and discover a strong resemblance to parallelized bidirectional search. Our findings advance understanding of the internal mechanisms underlying planning behavior in agents, which is important given the recent trend of emergent planning and reasoning capabilities in LLMs through RL
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。