arXiv:2603.24093cs.LGcs.AI2026-03被引 2

让大模型像人一样用外部经验+内部知识双重引导强化学习。

Towards Effective Experiential Learning: Dual Guidance for Utilization and Internalization

  • 构建经验库+内部知识双路径引导探索
  • 在多个推理任务上显著优于基线方法
  • 适合需要强逻辑推理的LLM训练场景

近期,强化学习(RL)已成为提升大语言模型(LLMs)能力的重要方法。特别是可验证奖励的强化学习(RLVR)在推理任务中展现出良好前景。然而,现有基于RL的训练仍仅是人类学习的粗略模拟:人类学习者会利用外部和内部经验指导探索,并逐步将有效轨迹内化为稳定知识。针对这一差距,我们提出统一框架DGO(Dual Guidance Optimization),通过外部经验和内部知识双路径提升训练效率。DGO首先从过往探索轨迹构建经验库,策略在经验库与模型内部知识的联合引导下进行探索;生成的新轨迹进一步用于更新经验库并优化模型参数,形成经验利用与内化的闭环。实验表明,DGO在多个推理任务上持续优于基线方法,证明更有效地利用与内化经验能显著提升推理能力。

原文摘要 · Abstract (English)

Recently, reinforcement learning~(RL) has become an important approach for improving the capabilities of large language models~(LLMs). In particular, reinforcement learning from verifiable rewards~(RLVR) has emerged as a promising paradigm for reasoning tasks. However, existing RL-based training still remains only a rough approximation to human learning. Human learners leverage both external and internal experience to guide exploration and gradually internalize useful trajectories into stable knowledge. Motivated by this gap, we ask: how can LLMs better utilize and internalize experience during RLVR training? To answer this question, we propose \textbf{D}ual \textbf{G}uidance \textbf{O}ptimization~(\textbf{DGO}), a unified framework that leverages \emph{external} and \emph{internal experience} to improve training effectiveness. Specifically, DGO first constructs an experience bank from previously explored trajectories. The policy then performs exploration under the joint guidance of the experience bank and the model's internal knowledge. The resulting trajectories are further used to refine the experience bank and optimize model parameters, forming a closed loop of experience utilization and internalization. Experiments show that DGO consistently outperforms baseline methods, suggesting that better utilization and internalization of experience lead to more effective reasoning.

强化学习推理能力大模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。