大模型在无奖励探索中也能积累知识,提升后续学习效率。
Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology
- 先无奖励探索,再引入奖励,促进知识组织
- 两阶段训练使模型性能优于全程奖励训练
- 为大模型突破奖励依赖提供心理学新思路
经典心理学中的潜在学习理论指出,生物体(如老鼠)可在无奖励情况下构建环境内部表征,一旦出现奖励即可快速适应。相比之下,认知科学认为当前的奖励学习仍过度依赖外部反馈,限制了灵活性与泛化能力。尽管大型语言模型(如OpenAI-o1和DeepSeek-R1)在推理能力上取得显著进展,但仍主要依赖以奖励为中心的强化学习范式。本研究首次揭示,大模型也表现出潜在学习动态:在初始的无奖励探索阶段,模型虽性能提升有限,但能不受奖励偏差约束地组织任务相关知识;当奖励引入后,性能进一步提升。采用两阶段探索策略的模型最终表现优于全程奖励训练的模型。我们通过跨多个模型家族和任务领域的广泛实验验证了这一现象,并提供了理论分析解释其机制。
原文摘要 · Abstract (English)
Latent learning, classically theorized by Tolman, shows that biological agents (e.g., rats) can acquire internal representations of their environment without rewards, enabling rapid adaptation once rewards are introduced. In contrast, from a cognitive science perspective, reward learning remains overly dependent on external feedback, limiting flexibility and generalization. Although recent advances in the reasoning capabilities of large language models (LLMs), such as OpenAI-o1 and DeepSeek-R1, mark a significant breakthrough, these models still rely primarily on reward-centric reinforcement learning paradigms. Whether and how the well-established phenomenon of latent learning in psychology can inform or emerge within LLMs' training remains largely unexplored. In this work, we present novel findings from our experiments that LLMs also exhibit the latent learning dynamics. During an initial phase of unrewarded exploration, LLMs display modest performance improvements, as this phase allows LLMs to organize task-relevant knowledge without being constrained by reward-driven biases, and performance is further enhanced once rewards are introduced. LLMs post-trained under this two-stage exploration regime ultimately achieve higher competence than those post-trained with reward-based reinforcement learning throughout. Beyond these empirical observations, we also provide theoretical analyses for our experiments explaining why unrewarded exploration yields performance gains, offering a mechanistic account of these dynamics. Specifically, we conducted extensive experiments across multiple model families and diverse task domains to establish the existence of the latent learning dynamics in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。